Patrick Le Callet

dblp:38/6666 · DBLP profile ↗
← Back
235ranked-venue papers
4as first author
103since 2021 · last 2026
0000-0002-2143-7063ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 221 · 2 first-author · 99 since 2021Human-computer interaction and ubiquitous computing · 25 · 11 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos: Methods and Results
Yixu Chen, Hai Wei, Pierre R. Lebreton, Patrick Le Callet, Alexander Kopte, Amritha Premkumar, Anna Meyer, Baojun Li, Changsheng Gao, Christian Herglotz, Christian Timmerer, Dandan Zhu 0001, Diwakara Reddy, Dong Liu 0002, Dounia Hammou, Guangtao Zhai, Hadi Amirpour, Hao Cheng 0015, Hichem Faraoun, Jonas Janzen, Krishna Srikar Durbha, Li Li 0040, Marc Windsheimer, MohammadAli Hamidi, Mykyta Skipenko, Paul Wawerek-Lopez, Pragyadipta Adhya, Prajit T. Rajendran, Rafal Mantiuk, Shien Ke, Sid Ahmed Fezza, Simon Deniffel, Wei Sun 0029, Weixia Zhang, Xiangguang Chen, Zuowei Cao, Minhao Tang, Xiaoyan Sun 0001, Xingwei Liu, Yeganeh Chatri, Yenan Xu
QoMEX5
2026 Assessing the impact of central and peripheral obstructions on visual behavior: Insights from gaze-contingent eye-tracking studies
abstract
Visual field loss, caused by conditions like glaucoma or macular degeneration, affects many people and impacts several life domains. This study contributes to the exploration of how people with visual field loss, such as central and peripheral scotomas, process visual stimuli in digital environments. Current visual attention models are based on experimental data obtained from individuals with normal vision, often overlooking those with limited vision. To address this issue, we compare gaze data from subjects viewing stimuli under two conditions: with and without visual-field masks of varying radii, with the main goal of understanding the role played by different components of vision in the overall grasping of visual information. We use metrics commonly employed for benchmarking saliency models as a means of assessing the similarity between data from each obstruction mask and the control stimuli, which could lead to the conclusion of whether foveal and peripheral vision contribute equally to natural vision or whether one of them stands out in information extraction. Novel saliency models could use this information to predict attention from visually-impaired individuals by possibly balancing these two sources of vision. Our results show a significantly higher similarity between control and central-scotoma saliency maps than between control and peripheral-scotoma data. Another statistical analysis shows no substantial learning effect or familiarity bias when participants revisit the same image under different conditions in the eye-tracking experiment. Finally, a difference-significance study reveals that different radii from central-scotoma conditions demonstrated no meaningful dispersion from each other.
Claudio M. S. Coutinho, Maria C. O. Faria, Alexandre Bruckert, Suiyi Ling, Matthieu Perreira Da Silva, Ronaldo F. Zampolo, Patrick Le Callet
Signal Process. Image Commun.7
2026 Temporal Consistency and Variation-Guided Spatio-Temporal Aggregation for Few-Shot Action Recognition
abstract
Few-shot Action Recognition (FSAR) aims to recognize novel actions from only a few labeled examples, posing challenges due to limited supervision and complex temporal dynamics. Existing methods often adopt a unified motion modeling strategy for both short- and long-term dynamics, overlooking the need to adapt motion pattern extraction to the specific temporal properties inherent to different timescales. This forces models to hedge against multi-scale relevance through exhaustive searches over temporal tuples, followed by heavy spatio-temporal fusion, which substantially increases parameters and computation and ultimately limits efficiency. To this end, we propose the efficient Temporal Consistency and Variation-Guided Spatio-Temporal Aggregation Network (TCV-STA), which comprises four key components: the Temporal Consistency Module (TCM), the Temporal Variation Module (TVM), the Spatio-Temporal Aggregation attention (STA), and the Shifted Window Temporal Attention (SWTA). The TCM captures stable motion patterns to suppress short-term perturbations and enhance temporal consistency for robust motion representation, while the TVM models dynamic motion patterns to highlight long-term variations that improve inter-class discriminability and facilitate intra-class alignment. Built upon these complementary motion cues, the STA selectively aggregates spatial and temporal representations under the guidance of the learned stable and dynamic motion patterns, avoiding global dense fusion. Finally, to address the limited receptive field and discontinuous modeling caused by frame grouping in TCM and TVM, we adapt a SWTA to capture longer-range temporal dependencies and ensure smooth transitions across subaction segments for few-shot action recognition. Experiments demonstrate that TCV-STA achieves competitive accuracy across four widely-used FSAR benchmarks while reducing parameters by up to 27.9% and computational cost by 21.3%, striking a favorable balance between accuracy and efficiency for deployment in resource-constrained scenarios.
Kaiwen Dong, Quanyi Li, Yanjing Sun, Xiao Yun, Yu Zhou 0009, Kévin Riou, Xiaofeng Hou, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.8
2026 Quality Assessment and Distortion-Aware Saliency Prediction for AI-Generated Omnidirectional Images
abstract
With the rapid advancement of Artificial Intelligence Generated Content (AIGC) techniques, AI generated images (AIGIs) have attracted widespread attention, among which AI generated omnidirectional images (AIGODIs) hold significant potential for Virtual Reality (VR) and Augmented Reality (AR) applications. AI generated omnidirectional images exhibit unique quality issues, however, research on the quality assessment and optimization of AI-generated omnidirectional images is still lacking. To this end, this work first studies the quality assessment and distortion-aware saliency prediction problems for AIGODIs, and further presents a corresponding optimization process. Specifically, we first establish a comprehensive database to reflecthumanfeedback for AI-generatedomnidirectionals, termed OHF2024, which includes both subjective quality ratings evaluated from three perspectives and distortion-aware salient regions. Based on the constructed OHF2024 database, we propose two models with shared encoders based on the BLIP-2 model to evaluate the human visual experience and predict distortion-aware saliency for AI-generated omnidirectional images, which are named as BLIP2OIQA and BLIP2OISal, respectively. Finally, based on the proposed models, we present an automatic optimization process that utilizes the predicted visual experience scores and distortion regions to further enhance the visual quality of an AI-generated omnidirectional image. Extensive experiments show that our BLIP2OIQA model and BLIP2OISal model achieve state-of-the-art (SOTA) results in the human visual experience evaluation task and the distortion-aware saliency prediction task for AI generated omnidirectional images, and can be effectively used in the optimization process. The database and codes will be released on https://github.com/IntMeGroup/AIGCOIQA to facilitate future research.
Huiyu Duan, Jing Liu 0002, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.8
2026 Deep Underwater Image Quality Assessment via Progressive Physics-Aware Multi-Prior Collaboration
abstract
Underwater image quality assessment (UIQA) is a critical research area, challenged by underwater environments such as wavelength-dependent light attenuation, scattering, and non-uniform illumination. Existing deep learning-based UIQA methods often address these degradations in isolation, neglecting their complex interplay with human perception and lacking explicit modeling of underwater optical phenomena. To address this, we propose PhysIQ-Net, a novel framework that integrates physics-driven principles with progressive multi-prior interaction modeling through three key innovations: First, introduce dual physics-based decomposition that separates images into Backscatter, Transmission, Reflectance, and Illuminance components to capture distinct degradation mechanisms; Second, propose prior-guided dynamic filtering that adapts convolutional kernels to image-specific content using physical priors; and Third, propose physic-informed Cross-Domain Feature Interaction that enables bidirectional collaboration between color-aware and structure-aware representations to model their perceptual inter-dependencies. Extensive experiments across multiple benchmark datasets demonstrate that PhysIQ-Net significantly outperforms existing methods, with ablation studies validating each component’s contribution, providing a robust solution for UIQA.
Zihan Zhou 0007, Jiaxue Lan, Yun Liang 0003, Jing Li 0026, Yong Xu 0007, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.7
2026 Multi-Dimensional Quality Assessment for Single-Image-to-3D Contents: Dataset and Model
abstract
The rapid advancement of AI generation technologies has led to the widespread use of AI-generated multimedia content, including images, videos, and 3D contents, across various applications. While significant progress has been made in quality evaluation for 2D content, evaluating the quality of 3D content synthesized from single image remains an underexplored problem. To bridge this gap, we introduce the first comprehensive subjective evaluation database tailored for assessing the quality of 3D content generated from single image. Our database, named AIGC-SI23DCQA, includes three distinct categories of input images, i.e., realistic images, AI-generated images, and computer graphic (CG) images, with 100 images in each category. Using five representative single-image-to-3D algorithms, we produce 1,500 3D contents and collect 94,500 annotations across three quality dimensions, including texture fidelity, shape accuracy, and overall quality. Based on the constructed database, we first benchmark and evaluate the performance of existing quality assessment methods revealing their limitations in addressing this novel task. Thus, we further propose a novel objective quality assessment method, termed I3DQA, for effective single-image-to-3D content quality assessment. Specifically, I3DQA first extracts the reference features from the source image, and the multi-modal features from the generated 3D content, including the projected video, patches, and large-multimodal model (LMM) features. These features are integrated through symmetric transformer blocks, enabling effective quality-related feature fusion and score prediction. Extensive experiments demonstrate the superior performance of our method and validate the effectiveness of its components. This work provides a foundational resource and a robust framework for advancing research in this emerging field, and our database and model are released at https://github.com/ZedFu/SI23DCQA.
Huiyu Duan, Jing Liu 0002, Yun Liu 0009, Xiaohong Liu 0001, Jia Wang 0004, Xiongkuo Min, Patrick Le Callet, Guangtao Zhai
IEEE Trans. Image Process.9
2025 Exploring The Potential of Vision-Language Models for Pure-Image and Text-Guided-Image Saliency Prediction
abstract
We introduce VLSal, a saliency prediction framework that leverages Vision-Language Models (VLMs) to unify pure-image and text-guided-image saliency prediction tasks and achieve high performance in both. We extract visual features from the visual encoder and retrieve the corresponding visual token features from the language decoder, which serves as a natural feature fusion mechanism. These features are then processed through a U-Net-based saliency decoder to generate accurate saliency maps. To efficiently adapt the large-scale pretrained model, we apply Low-Rank Adaptation (LoRA) finetuning, reducing computational costs while preserving performance. Extensive experiments on benchmark datasets, including SALICON, MIT1003, and TIS, demonstrate that VLSal outperforms existing methods in both pure-image and text-guided-image saliency prediction.
Huiyu Duan, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ICIP5
2025 InpaintFormer: Prompt-guided High-Quality Face Inpainting with Mask-Aware Self-Attention
abstract
Face image inpainting, especially with user-controllable customization, aims to restore degraded facial regions while adhering to user-provided instructions. Traditional inpainting methods often focus solely on restoring visual fidelity, lacking the ability to incorporate user prompts or semantic guidance. In this work, we present InpaintFormer, a novel framework for user-controlled face image inpainting guided by textual prompts. Specifically, we propose a Prompt-guided Feature Modulation (PGFM) module to align visual features with user instructions by utilizing a pre-trained CLIP model to extract text and image embeddings. These embeddings are fused to modulate the encoded image features, ensuring semantic consistency with the prompt. Additionally, a Degradation Mask Predictor (DMP) is introduced to identify degraded regions requiring inpainting, while a Mask-Aware Self-Attention (MASA) mechanism within the Transformer refines the inpainting process by selectively attending to non-degraded regions for generating realistic results. By combining PGFM, DMP, and MASA, InpaintFormer enables controllable face image inpainting with high fidelity and semantic alignment. Extensive experiments demonstrate that InpaintFormer outperforms state-of-the-art inpainting methods in terms of controllability and naturalness.
Zhouhao Ouyang, Yan Huang 0031, Si Wu 0002, Yong Xu 0007, Patrick Le Callet, Dapeng Oliver Wu
ICME7
2025 Rethinking 3D Robotic Perception: Elastic Voxel Representation with Splatting Distillation
abstract
Language-guided robotic manipulation is advancing rapidly with Vision-Language-Action (VLA) models, yet faces fundamental challenges in 3D perception. This paper addresses two critical challenges: the scale elasticity requirement for simultaneously processing coarse environmental context and fine manipulation details, and the scarcity of action-annotated training data. We present Splat-Actor, a novel robotic manipulation framework that introduces two key innovations. First, we develop an elastic voxel encoder that combines multi-scale processing with selective tokenization, enabling efficient 3D spatial reasoning while adaptively focusing on informative regions. Second, we propose a depth-constrained feature distillation framework that leverages Gaussian Splatting to bridge 2D and 3D representations, transferring rich semantic features from pre-trained vision models to enhance 3D understanding. Extensive experiments across 10 manipulation tasks with 166 variations demonstrate that Splat-Actor achieves a 6.8% improvement over state-of-the-art methods while maintaining the computational efficiency.
Shaohui Pan, Yong Xu 0007, Ruotao Xu, Zihan Zhou 0007, Si Wu 0002, Zhu Liang Yu, Patrick Le Callet
ICME7
2025 Text to Trajectory: Enhancing and Evaluating LLMs for Embodied Task Planning
abstract
The increasing demand for effective human-machine interaction highlights the importance of integrating natural language processing with robotics technology. This paper addresses the challenges of using Large Language Models (LLMs) for embodied task planning in complex environments. We propose a comprehensive framework that combines environmental-aware LLM fine-tuning with a novel Stepwise Beam Search (SBS) strategy. In conjunction with the environmentally enhanced LLM, the SBS strategy facilitates comprehensive exploration of both token-level and step-level search spaces, overcoming the limitations of conventional greedy search methods. Additionally, to evaluate the effectiveness of embodied task planning, we introduce the Trajectory Match Score (TMS), a robust evaluation metric that leverages state-based simulation to assess plan success. Through extensive experiments on standard benchmarks, our framework demonstrates substantial improvements in both plan generation quality and task success rates, advancing the state-of-the-art in embodied task planning.
Yihan Tang, Yong Xu 0007, Ruotao Xu, Yan Huang 0031, Si Wu 0002, Patrick Le Callet
ICME6
2025 SemanticLoom: Category-aware Dynamic Fusion for Multi-class Few-shot Image Synthesis
abstract
Few-shot text-to-image (T2I) generation seeks to efficiently integrate new semantics into existing pre-trained models while preserving their capacity to generate diverse, high-quality images. However, existing methods often suffer from inefficiency and poor scalability due to the need for separate training processes for each new concept. These challenges hinder their practical application in multi-class few-shot scenarios. To overcome these issues, we propose SemanticLoom that dynamically incorporates novel concepts into pre-trained diffusion models through category-aware dynamic feature fusion. Our approach introduces a lightweight semantic expander that captures fine-grained semantics, guided by learnable identifier to ensure precise semantic integration. By dynamically adjusting feature fusion coefficients based on category diversity and training progress, our method harmonizes the integration of new semantic features with the original model’s capabilities, ensuring consistency and generalization. Experiments demonstrate that our method successfully integrates new semantics without compromising the generative diversity and versatility of the pre-trained model.
Yan Huang 0031, Si Wu 0002, Yong Xu 0007, Patrick Le Callet
ICME7
2025 Adaptive Illumination Transfer Network for Shadow Removal
abstract
Shadow removal aims to harmonize illumination between shadow and non-shadow regions. However, existing methods often struggle to achieve this goal due to inadequate modeling of illumination relationships between these two regions. Moreover, the prevalent reliance on binary shadow masks hinders their capability to address non-uniform shadows. To address these limitations, we propose an adaptive illumination transfer network (AITNet), which incorporates two complementary shadow enhancement strategies. First, a global illumination transfer strategy is designed to model the illumination relationship between shadow and non-shadow regions, enabling the holistic enhancement of shadow regions. Second, an illumination-adaptive strategy is developed to estimate an illumination degradation map, which guides the adaptive enhancement of shadow regions. Furthermore, to preserve the original structural details, the enhancement process is applied exclusively to the illumination map obtained after Retinex decomposition. Extensive experiments have demonstrated the superiority of our method over existing approaches on public datasets.
Si Wu 0002, Yong Xu 0007, Yan Huang 0031, Patrick Le Callet
ICME5
2025 HarmonyIQA: Pioneering Benchmark and Model for Image Harmonization Quality Assessment
abstract
Image composition involves extracting a foreground object from one image and pasting it into another image through Image harmonization algorithms (IHAs), which aim to adjust the appearance of the foreground object to better match the background. Existing image quality assessment (IQA) methods may fail to align with human visual preference on image harmonization due to the insensitivity to minor color or light inconsistency. To address the issue and facilitate the advancement of IHAs, we introduce the first Image Quality Assessment Database for image Harmony evaluation (HarmonyIQAD), which consists of 1,350 harmonized images generated by 9 different IHAs, and the corresponding human visual preference scores. Based on this database, we propose a Harmony Image Quality Assessment (HarmonyIQA), to predict human visual preference for harmonized images. Extensive experiments show that HarmonyIQA achieves state-of-the-art performance on human visual preference evaluation for harmonized images, and also achieves competing results on traditional IQA tasks. Furthermore, cross-dataset evaluation also shows that HarmonyIQA exhibits better generalization ability than self-supervised learning-based IQA methods. The dataset and code are available at https://github.com/IntMeGroup/HarmonyIQA.
Zitong Xu, Huiyu Duan, Guangji Ma, Qingbo Wu 0001, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ICME9
2025 ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos
abstract
With the rapid development of eXtended Reality (XR), egocentric spatial shooting and display technologies have further enhanced immersion and engagement for users, delivering more captivating and interactive experiences. Assessing the quality of experience (QoE) of egocentric spatial videos is crucial to ensure a high-quality viewing experience. However, the corresponding research is still lacking. In this paper, we use the concept of embodied experience to highlight this more immersive experience and study the new problem, i.e., embodied perceptual quality assessment for egocentric spatial videos. Specifically, we introduce the first Egocentric Spatial Video Quality Assessment Database (ESVQAD), which comprises 600 egocentric spatial videos captured using the Apple Vision Pro and their corresponding mean opinion scores (MOSs). Furthermore, we propose a novel multi-dimensional binocular feature fusion model, termed ESVQAnet, which integrates binocular spatial, motion, and semantic features to predict the overall perceptual quality. Experimental results demonstrate the ESVQAnet significantly outperforms 16 state-of-the-art VQA models on the embodied perceptual quality assessment task, and exhibits strong generalization capability on traditional VQA tasks. The database and code are available at https://github.com/IntMeGroup/ESVQA.
Xilei Zhu, Huiyu Duan, Yucheng Zhu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ICME7
2025 Evaluating Perceptual Color Preferences in Smartphone Photography: Dataset and Challenges
abstract
International audience
Zhihua Wang 0002, Weixia Zhang, Wei Zhou 0021, Xiaohong Liu 0001, Guangtao Zhai, Patrick Le Callet
ACM Multimedia6
2025 Omni2: Unifying Omnidirectional Image Generation and Editing in an Omni Model
abstract
360° omnidirectional images (ODIs) have gained considerable attention recently, and are widely used in various virtual reality (VR) and augmented reality (AR) applications. However, capturing such images is expensive and requires specialized equipment, making ODI synthesis increasingly important. While common 2D image generation and editing methods are rapidly advancing, these models struggle to deliver satisfactory results when generating or editing ODIs due to the unique format and broad 360° Field-of-View (FoV) of ODIs. To bridge this gap, we construct Any2Omni , the first comprehensive ODI generation-editing dataset comprises 60,000+ training data covering diverse input conditions and up to 9 ODI generation and editing tasks. Built upon Any2Omni, we propose an Omni model for Omni-directional image generation and editing ( Omni 2), with the capability of handling various ODI generation and editing tasks under diverse input conditions using one model. Extensive experiments demonstrate the superiority and effectiveness of the proposed Omni2 model for both the ODI generation and editing tasks. Both the Any2Omni dataset and the Omni2 model are publicly available at: https://github.com/IntMeGroup/Omni2.
Huiyu Duan, Yucheng Zhu, Xiaohong Liu 0001, Lu Liu 0005, Zitong Xu, Guangji Ma, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ACM Multimedia10
2025 Introducing VMAF-AC, A Visual Quality Metric For Asymmetric Video Coding: Use Case on Sport Video Content
Shivam Bhardwaj, Pierre R. Lebreton, Tushar Shinde, Patrick Le Callet
PCS4
2025 Study on content-dependency of acceptability/annoyance (AccAnn) scale in User-Generated Content (UGC) videos
abstract
International audience
Pierre R. Lebreton, Patrick Le Callet, Neil Birkbeck, Yilin Wang 0001, Zeina Sinno, Balu Adsumilli
PCS2
2025 Data Augmentation for QoL-Centered Functional Vision Research: Synthetic Human Behavior Generation in Virtual Reality
abstract
Functional vision assessment is essential for understanding the Quality of Life (QoL) of individuals with visual impairments. The Multi-Luminance Mobility Test (MLMT) is a promising Orientation and Mobility (O&M) method that provides objective functional vision evaluation. However, the current test primarily relies on a statistical factor-based scoring system. Additionally, the scarcity of behavioral data, collection difficulties, and privacy concerns hinder the development of more detailed, behavior-based evaluation metrics. To address these challenges, we propose a data augmentation approach leveraging a Virtual Reality (VR)-based O&M test protocol combined with diffusion policy-based models to generate synthetic behavioral data. In this study, we adapted a transformer-based diffusion policy to generate multi-dimensional motion sequences under varying luminance conditions from VR-based O&M protocols. Quantitative evaluations demonstrate that the synthetic data effectively captures the relationship between luminance and motion for luminance levels seen during training. The zero-shot generalization ability of the policy is also explored. Our findings suggest that diffusion policy-generated synthetic data can enhance functional vision research by addressing data scarcity and supporting the development of behavior-based assessment metrics. The code is available at https://gitlab.univ-nantes.fr/E21A837H/diffusionpolicyvr_motiongeneration.git.
Kévin Riou, Alexandre Bruckert, Patrick Le Callet
QoMEX4
2025 An HMM-Based Behavior Analysis Approach for QoE in VR : A Case Study of QoL Assessment via Orientation and Mobility Test
abstract
Virtual Reality (VR), as a leading form of immersive media, has rapidly expanded into various application domains, raising the need for standardized methods to evaluate its Quality of Experience (QoE). An ongoing recommendation, ITU-T P.IXC, titled "Interactive test methods for subjective assessment of XR communications", is currently under joint development by the Video Quality Experts Group Immersive Media Group (VQEG-IMG) and ITU-T Study Group 12 (SG12). One major challenge identified in this context is the insufficient attention paid to user behaviors during tasks, despite the increasing availability of behavioral data. Although such data are often collected, they are rarely analyzed in depth. In this paper, we proposed a Hidden Markov Model (HMM)-based behavior analysis approach. Using a VR-based Orientation and Mobility (O&M) test designed for Quality of Life (QoL) assessment–one of the promising applications of VR–we leveraged the collected behavioral data to model behavior patterns. Our results demonstrate that the proposed approach can effectively extract behavior-aware features, highlighting the potential of integrating behavior analysis into QoL assessment frameworks. We believe this work provides valuable insights into incorporating behavioral metrics into immersive media QoE evaluation and contributes to the development of ITU-T P.IXC.
Alexandre Bruckert, Patrick Le Callet
VCIP3
2025 Is there a relationship between Mean Opinion Score (MOS) and Just Noticeable Difference (JND)?
abstract
Evaluating perceived video quality is essential for ensuring high Quality of Experience (QoE) in modern streaming applications. While existing subjective datasets and Video Quality Metrics (VQMs) cover a broad quality range, many practical use cases—especially for premium users—focus on high-quality scenarios requiring finer granularity. Just Noticeable Difference (JND) has emerged as a key concept for modeling perceptual thresholds in these high-end regions and plays an important role in perceptual bitrate ladder construction. However, the relationship between JND and the more widely used Mean Opinion Score (MOS) remains unclear. In this paper, we conduct a Degradation Category Rating (DCR) subjective study based on an existing JND dataset to examine how MOS corresponds to the 75% Satisfied User Ratio (SUR) points of the 1stand 2ndJNDs. We find that while MOS values at JND points generally align with theoretical expectations (e.g., 4.75 for the 75% SUR of the 1stJND), the reverse mapping—from MOS to JND—is ambiguous due to overlapping confidence intervals across PVS indices. Statistical significance analysis further shows that DCR studies with limited participants may not detect meaningful differences between reference and JND videos.
Hadi Amirpour, Wei Zhou 0021, Patrick Le Callet
VCIP4
2025 ERD: Encoder-Residual-Decoder Neural Network for Underwater Image Enhancement
abstract
In underwater environments, the absorption and scattering of light often result in various types of degradation in captured images, including color cast, low contrast, low brightness, and blurriness. These undesirable effects pose significant challenges for both underwater photography and downstream tasks such as object detection, recognition, and navigation. To address these challenges, we propose a novel end-to-end underwater image enhancement (UIE) network via the multistage and mixed attention mechanism and a residual-based feature refinement module, called ERD. Specifically, our network includes an encoder stage for extracting features from input underwater images with channel, spatial, and patch attention modules to emphasize degraded channels and regions for restoration; a residual stage for further purification of informative features through sufficient feature learning; and a decoder stage for effective image reconstruction. Inspired by visual perception mechanism, we design the frequency domain loss and edge details loss to retain more high-frequency information and object details while ensuring that the enhanced image approximates the reference image in terms of color tone while preserving content and structure. To comprehensively evaluate our proposed UIE model, we also curated three additional underwater image datasets through online collection and generation using Cycle-GAN. Rigorous experiments conducted on a total of eight underwater image datasets demonstrate that the proposed ERD model outperforms state-of-the-art methods in enhancing both real-world and generated underwater images. Our code and datasets are available athttps://github.com/fansuregrin/ERD.
Jingchao Cao, Wangzhen Peng, Yutao Liu 0002, Junyu Dong, Patrick Le Callet, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.5
2025 Detail-Preserving Diffusion Models for Low-Light Image Enhancement
abstract
Existing diffusion models for low-light image enhancement typically incrementally remove noise introduced during the forward diffusion process using a denoising loss, with the process being conditioned on input low-light images. While these models demonstrate remarkable abilities in generating realistic high-frequency details, they often struggle to restore fine details that are faithful to the input. To address this, we present a novel detail-preserving diffusion model for realistic and faithful low-light image enhancement. Our approach integrates a size-agnostic diffusion process with a reverse process reconstruction loss, significantly enhancing the fidelity of enhanced images to their low-light counterparts and enabling more accurate recovery of fine details. To ensure the preservation of region- and content-aware details, we employ an efficient noise estimation network with a simplified channel-spatial attention mechanism. Additionally, we propose a multiscale ensemble scheme to maintain detail fidelity across diverse illumination regions. Comprehensive experiments on eight benchmark datasets demonstrate that our method achieves state-of-the-art results compared to over twenty existing methods in terms of both perceptual quality (LPIPS) and distortion metrics (PSNR and SSIM). The code is available at:https://github.com/CSYanH/DePDiff.
Yan Huang 0031, Xiaoshan Liao, Jinxiu Liang, Boxin Shi, Yong Xu 0007, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.6
2025 Subjective and Objective Audio-Visual Quality Assessment for Omnidirectional Videos
abstract
Virtual Reality (VR) has attracted widespread attention in recent years due to its capability to create immersive experiences by presenting multi-modal information to users. Omnidirectional videos (ODVs), as a prominent component of VR content, are essential across diverse applications. This necessitates service providers to monitor and optimize the quality of ODVs throughout the filming, encoding, decoding, and transmission stages to ensure a high-quality viewing experience. However, most existing Quality of Experience (QoE) studies for ODVs only focus on the visual quality, while overlooking the impact of the audio modality on perceptual quality. This paper presents a comprehensive study of omnidirectional audio-visual quality assessment (OD-AVQA) from both subjective and objective perspectives. Specifically, we first establish a large-scale audio-visual quality assessment database for ODVs named OAVQAD+, which includes 625 distorted omnidirectional audio-visual sequences derived from 25 pristine ODVs, and the corresponding collected mean opinion scores (MOSs) for the QoE of these ODVs. This contributes to the largest database for assessing the audio-visual quality of ODVs. To advance the fields of objective OD-AVQA, we construct a benchmark that includes three types of benchmark models. Type I and Type II models integrate well-known video quality assessment (VQA) and audio quality assessment (AQA) methods using support vector regression (SVR) and multi-layer perceptron (MLP), respectively, while Type III consists of AVQA models specifically designed for traditional 2D audio-visual sequences. We also propose a novel Omnidirectional Audio-Visual quality assessment Network (OmniAVNet) that integrates quality-aware audio, visual, and motion features to predict overall audio-visual quality for ODVs effectively, which supports both full-reference (FR) and no-reference (NR) assessment. Extensive experimental results demonstrate that OmniAVNet outperforms the aforementioned benchmark OD-AVQA models on two OD-AVQA databases, and shows great performance on one omnidirectional VQA database. The database and code are available at https://github.com/IntMeGroup/OmniAVNet.
Xilei Zhu, Huiyu Duan, Yuqin Cao, Yucheng Zhu, Jing Liu 0002, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
IEEE Trans. Image Process.9
2025 VQM4HAS: A Real-Time Quality Metric for HEVC Videos in HTTP Adaptive Streaming
abstract
In HTTP Adaptive Streaming (HAS), a video is encoded at various bitrate-resolution pairs, collectively known as the bitrate ladder, allowing users to select the most suitable representation based on their network conditions. Optimizing this set of pairs to enhance the Quality of Experience (QoE) requires accurately measuring the quality of these representations. VMAF and ITU-T's P.1204.3 are highly reliable metrics for assessing the quality of representations in HAS. However, in practice, using these metrics for optimization is often impractical for live streaming applications due to their high computational costs and the large number of bitrate-resolution pairs in the bitrate ladder that need to be evaluated. To address their high complexity, our paper introduces a new method calledVQM4HAS, which extractslow-complexityfeatures, including ($i$) video complexity features, ($ii$) frame-level encoding statistics logged during the encoding process, and ($iii$) lightweight video quality metrics. These extracted features are then fed into a regression model to predict VMAF or P.1204.3. TheVQM4HASmodel is designed to operate on a per bitrate-resolution pair, per-resolution, and cross-representation basis, optimizing quality predictions across different scenarios. Our experimental results demonstrate thatVQM4HASachieves a high correlation with VMAF and P.1204.3, with Pearson correlation coefficients (PCC) ranging from 0.95 to 0.96 for VMAF and 0.97 to 0.99 for P.1204.3, depending on the resolution. Despite achieving a high correlation with VMAF and P.1204.3,VQM4HASexhibits significantly less complexity than both metrics, with 98% and 99% less complexity for VMAF and P.1204.3, respectively, making it suitable for live streaming scenarios. We also conduct a feature importance analysis to further reduce the complexity of the proposed method. Furthermore, we evaluate the effectiveness of our method by using it to predict subjective quality scores. The results show thatVQM4HASachieves a higher correlation with subjective scores at various resolutions despite its minimal complexity. The source code is available athttps://github.com/cd-athena/VQM4HAS.
Hadi Amirpour, Wei Zhou 0021, Patrick Le Callet, Christian Timmerer
IEEE Trans. Multim.4
2025 Enhancing CNN-Based Blind Image Quality Assessment via Deep Cross-Layer Pattern Encoding
abstract
Evaluating image quality without reference images, known as blind image quality assessment (BIQA), is crucial for image communication. Recently, convolutional neural networks (CNNs) have emerged as a prominent BIQA approach due to their feature learning power. Usually, both high-level semantic information and low-level details significantly impact perceived visual quality. However, most existing CNN-based methods focus on high-level semantic information via aggregating features on top of the last convolutional layer into a global descriptor, neglecting the importance of shallow, low-level cues. To address this limitation, this paper proposes a novel approach that exploits local encoding and histogram-based pyramid pooling on crosslayer features produced by a CNN, achieving a joint local and global analysis. Specifically, we introduce a cross-layer pattern encoding model that characterizes features generated along convolutional layers via a soft histogram of local 3D binary patterns. This leads to a highly informative yet compact descriptor for score regression. By building this module into a ResNet backbone, we present an effective BIQA model demonstrating state-ofthe-art performance in extensive experiments on synthetic and authentic datasets.
Zihan Zhou 0007, Yong Xu 0007, Yuhui Quan, Yun Liang 0003, Jing Li 0026, Patrick Le Callet
IEEE Trans. Multim.6
2025 Interactions Between Vibroacoustic Discomfort and Visual Stimuli: Comparison of Real, 3D and 360 Environments
abstract
The building industry and the design of interior environments are increasingly focusing on the user experience, incorporating sensory analysis to reconsider how office environments can be optimized. New immersive technologies offer significant opportunities for sensory science, enhancing our understanding of human perception and enabling the collection of multi-sensory data under controlled laboratory conditions. While the potential of Virtual Reality (VR) for these types of studies is well recognized, certain limitations still need to be addressed, including the lack of standardized research practices and the challenge of ensuring the simulated environment closely mirrors the real world. In this study, we compare 360° and 3D formats, to real-life settings in order to determine which format offers greater ecological validity for visual perception and immersion. Additionally, we examine the effects of vibroacoustic stimuli with different levels of intensity on perception and cognition of 30 participants. Subjective, physiological and cognitive data was collected throughout the test to tackle the participant's experience. This preliminary study introduces an immersive methodology that leverages advanced techniques to gain deeper insights into multisensory user experience in VR, marking a significant step forward in the optimization of VR for building evaluation.
Charlotte Scarpa, Toinon Vigier, Gwénaëlle Haese, Patrick Le Callet
IEEE Trans. Vis. Comput. Graph.4
2025 ESIQA: Perceptual Quality Assessment of Vision-Pro-based Egocentric Spatial Images
abstract
With the development of eXtended Reality (XR), photo capturing and display technology based on head-mounted displays (HMDs) have experienced significant advancements and gained considerable attention. Egocentric spatial images and videos are emerging as a compelling form of stereoscopic XR content. The assessment for the Quality of Experience (QoE) of XR content is important to ensure a high-quality viewing experience. Different from traditional 2D images, egocentric spatial images present challenges for perceptual quality assessment due to their special shooting, processing methods, and stereoscopic characteristics. However, the corresponding image quality assessment (IQA) research for egocentric spatial images is still lacking. In this paper, we establish the Egocentric Spatial Images Quality Assessment Database (ESIQAD), the first IQA database dedicated for egocentric spatial images as far as we know. Our ESIQAD includes 500 egocentric spatial images and the corresponding mean opinion scores (MOSs) under three display modes, including 2D display, 3D-window display, and 3D-immersive display. Based on our ESIQAD, we propose a novel mamba2-based multi-stage feature fusion model, termed ESIQAnet, which predicts the perceptual quality of egocentric spatial images under the three display modes. Specifically, we first extract features from multiple visual state space duality (VSSD) blocks, then apply cross attention to fuse binocular view information and use transposed attention to further refine the features. The multi-stage features are finally concatenated and fed into a quality regression network to predict the quality score. Extensive experimental results demonstrate that the ESIQAnet outperforms 22 state-of-the-art IQA models on the ESIQAD under all three display modes. The database and code are available at https://github.com/IntMeGroup/ESIQA.
Xilei Zhu, Huiyu Duan, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
IEEE Trans. Vis. Comput. Graph.6
2024 Behavioral Recognition of Skeletal Data Based on Targeted Dual Fusion Strategy
abstract
The deployment of multi-stream fusion strategy on behavioral recognition from skeletal data can extract complementary features from different information streams and improve the recognition accuracy, but suffers from high model complexity and a large number of parameters. Besides, existing multi-stream methods using a fixed adjacency matrix homogenizes the model’s discrimination process across diverse actions, causing reduction of the actual lift for the multi-stream model. Finally, attention mechanisms are commonly applied to the multi-dimensional features, including spatial, temporal and channel dimensions. But their attention scores are typically fused in a concatenated manner, leading to the ignorance of the interrelation between joints in complex actions. To alleviate these issues, the Front-Rear dual Fusion Graph Convolutional Network (FRF-GCN) is proposed to provide a lightweight model based on skeletal data. Targeted adjacency matrices are also designed for different front fusion streams, allowing the model to focus on actions of varying magnitudes. Simultaneously, the mechanism of Spatial-Temporal-Channel Parallel Attention (STC-P), which processes attention in parallel and places greater emphasis on useful information, is proposed to further improve model’s performance. FRF-GCN demonstrates significant competitiveness compared to the current state-of-the-art methods on the NTU RGB+D, NTU RGB+D 120 and Kinetics-Skeleton 400 datasets. Our code is available at: https://github.com/sunbeam-kkt/FRF-GCN-master.
Xiao Yun, Kévin Riou, Kaiwen Dong, Yanjing Sun, Song Li 0001, Kévin Subrin, Patrick Le Callet
AAAI8
2024 Exploring Bitrate Costs for Enhanced User Satisfaction: A Just Noticeable Difference (JND) Perspective
abstract
The evolving landscape of the delivery of multimedia content requires a deep understanding of how the design of the bitrate ladder (for HTTP adaptive streaming) influences cost and quality. This paper explores the use of Just Noticeable Differences (JND) to select bitrate-resolution pairs for constructing a bitrate ladder with respect to the proportion of satisfied user ratio (SUR). To expand the investigation to various codecs, first, a method is explained that transfers the JND points obtained through subjective testing from one codec (e.g., , AVC) to other codecs (e.g., , HEVC, VVC). This approach helps avoid the additional costs associated with conducting subjective tests to obtain JND points for a wide range of different codecs. To achieve this objective, we investigate the codec-agnostic nature of various video quality metrics, followed by the transfer of JND between two codecs, taking into account the most suitable codec-agnostic video quality metric. Secondly, we delve into the analysis of the bitrate cost of a given bitrate ladder from a JND perspective, i.e., , as a function of the SUR. Among others, our experimental results demonstrate that increasing SUR leads to an exponential increase in bitrate. For example, to raise the SUR from 75% to 90%, it is necessary to double the video bitrate.
Hadi Amirpour, Raimund Schatz, Patrick Le Callet, Christian Timmerer
DCC4
2024 A Real-Time Video Quality Metric for HTTP Adaptive Streaming
abstract
In HTTP Adaptive Streaming (HAS), a video is encoded at multiple bitrate-resolution pairs, referred to as representations, which enables users to choose the most suitable representation based on their network connection. To optimize the set of bitrate-resolution pairs and improve the Quality of Experience (QoE) for users, it is of utmost importance to measure the quality of the representations. VMAF is a highly reliable metric used in HAS to assess the quality of representations. However, in practice, using it for optimization can be a very time-consuming process, and it is infeasible for live streaming applications. To tackle its high complexity, our paper introduces a new method called VQM4HAS, which extracts low-complexity features, including (i) video complexity features, (ii) bitstream features logged during the encoding process, and (iii) basic video quality metrics. These extracted features are then fed into a regression model to predict VMAF. Our experimental results demonstrate that VQM4HAS achieves a high Pearson Correlation Coefficient (PCC) with VMAF, ranging from 0.95 to 0.96 depending on the resolution. However, it exhibits significantly lower complexity, making it suitable for live streaming scenarios.
Hadi Amirpour, Patrick Le Callet, Christian Timmerer
ICASSP3
2024 Comparison of Conditions for Omnidirectional Video with Spatial Audio in Terms of Subjective Quality and Impacts on Objective Metrics Resolving Power
abstract
Omnidirectional media formats, particularly 360° videos with spatial audio, provide new immersive experiences and introduce a novel dimension to content consumption.We explore the relationship between subjective data quality and metric performance evaluation in the context of Omnidirectional videos with spatial audio. While methodologies for 360° video quality assessment have been standardized and well-documented, previous efforts primarily focus on video with limited audio conditions, e.g., mono/stereo rendering. Moreover, the experimental test setup and subjective test methodologies impact data quality and the ability to use these data for objective quality metrics performance evaluation. Such a problem is key in the industry and the standardization activities, as codecs and quality models must be compared. Hence, the requirements on the ground truth data quality have to be clarified to allow proper conclusions.In this paper, we compare two setups and three test methodologies to study how experiment discriminability changes with conditions and participant number. Then, we show how discriminability impacts the resolving power of quality metrics. We show that higher-performing metrics require higher-quality data to reveal their full potential. In doing so, we put into relation the experimental cost, data quality, and resolving power.
Andreas Pastor, Pierre R. Lebreton, Toinon Vigier, Patrick Le Callet
ICASSP4
2024 Comparison of Crowdsourcing And Laboratory Settings for Subjective Assessment of Video Quality and Acceptability & Annoyance
abstract
User satisfaction is significantly influenced by their expectations of video quality. Even when users are presented with identical video stimuli, the Quality of Experience (QoE) can vary based on the context. The acceptability and annoyance paradigm serves as a tool to understand this relationship by measuring QoE as a function of user expectations and video quality. Traditionally, subjective experiments assessing QoE have been conducted in controlled laboratory settings. While the extension of traditional video quality experiments to crowdsourcing settings is well-explored, the impact of crowdsourcing on QoE studies has not been thoroughly examined. This study explore the potential use of crowdsourcing platforms for acceptability & annoyance experiments. To this end, video quality and acceptability & annoyance experiments were conducted in both laboratory and crowdsourcing settings. The findings reveal a more linear relationship between video quality and QoE in crowdsourcing settings. Subjects in crowdsourcing settings tend to have higher expectations of video quality, resulting in a slight increase in acceptability & annoyance thresholds compared to laboratory experiments. Analyses suggest that extending acceptability & annoyance experiments to crowdsourcing is not as straightforward as extending traditional video quality experiments. In crowdsourcing settings, priming subject expectations with instructions is not as effective as it is in laboratory conditions.
Ali Ak, Abhishek Gera, Denise Noyes, Hassene Tmar, Ioannis Katsavounidis, Patrick Le Callet
ICIP6
2024 A Toolkit to Benchmark Point Cloud Quality Metrics with Multi-Track Evaluation Criteria
abstract
Point clouds (PCs) gained popularity as a representation for 3D objects and scenes and are widely used in numerous applications in augmented and virtual reality domains. Concurrently, quality assessment of PCs became even more relevant to improve various aspects of these imaging pipelines. To stimulate further growth and interest in point cloud quality assessment (PCQA), we created a large-scale PCQA dataset (called “BASICS”) which provides the research community with a relevant and challenging dataset to develop reliable objective quality metrics, and we organized the PCVQA grand challenge at ICIP 2023. In this paper, we provide a track-based evaluation methodology for benchmarking visual quality metrics, mirroring the PCVQA grand challenge evaluation scenarios designed to mimic real-life applications. Furthermore, we provide a state-of-the-art benchmark for the point cloud quality metrics. The track-based benchmarking approach shows that there is room for improvement in certain research directions, drawing attention to open problems in the PCQA domain.
Ali Ak, Emin Zerman, Maurice Quach, Aladine Chetouani, Giuseppe Valenzise, Patrick Le Callet
ICIP6
2024 A Dataset for Understanding Open UGC Video Datasets
abstract
User Generated Content (UGC) video streaming is a major application on the Internet. Even small bitrate savings can have large network impacts at this scale. In order to achieve improvements without sacrificing experience, the quality of UGC videos needs to be better understood. In recent years video quality evaluation models designed for the evaluation of UGC videos have received a lot of attention. However, considering that these models are learning-based models, they heavily depend on the training data that has been used. In this paper, a new dataset is introduced that allows studying the differences in characteristics between existing UGC video datasets. It reveals the range of quality that was covered by existing UGC video datasets, and the implication of these quality ranges on training and validation performance of UGC video quality prediction models. Furthermore, this work demonstrates that dataset alignment enables existing UGC models to achieve higher performance. This alignment dataset can be found openly available on Zenodo (https://zenodo.org/doi/10.5281/zenodo.12155934).
Pierre R. Lebreton, Patrick Le Callet, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli
ICIP2
2024 Evaluating 3D Human Pose Estimation in Occluded Multi-Sensor Scenarios: Dataset and Annotation Approach
abstract
Obtaining ground truth annotations for 3D pose estimation (3D HPE) typically depends on motion capture equipment (Mocap), which is not only expensive but impractical for widespread deployment. In contrast, triangulation can reconstruct 3D poses solely from multi-view 2D poses with known camera parameters, eliminating the need for Mocap. However, inherent noise in 2D pose predictions introduces uncertainties, compromising the reliability of the results. To obtain more reliable annotations with noisy input, we introduce an annotation approach for the 3D HPE task, driven by prior knowledge of the skeletal configuration. We split our approach into two steps: first a parametric model is designed to enhance confidence predictions. Then, a differentiable weighted triangulation is employed to estimate the 3D pose in world space, leveraging the predicted confidence scores as weights. The pipeline is trained using a bone length loss. Moreover, we collect a multi-view dataset for 3D HPE and annotate it using our proposed annotation tool. This dataset is characterized by more construction scenarios, including heavier occlusion cases, diverse viewing directions, and the integration of various optical sensors, setting it apart from existing datasets. Experiments on both our dataset and Human3.6M demonstrate the effectiveness of our method.
Kévin Riou, Kaiwen Dong, Kévin Subrin, Patrick Le Callet, Yanjing Sun
ICIP5
2024 AIGCOIQA2024: Perceptual Quality Assessment of AI Generated Omnidirectional Images
abstract
[?]In recent years, the rapid advancement of Artificial Intelligence Generated Content (AIGC) has attracted widespread attention. Among the AIGC, AI generated omnidirectional images hold significant potential for Virtual Reality (VR) and Augmented Reality (AR) applications, hence omnidirectional AIGC techniques have also been widely studied. AI-generated omnidirectional images exhibit unique distortions compared to natural omnidirectional images, however, there is no dedicated Image Quality Assessment (IQA) criteria for assessing them. This study addresses this gap by establishing a large-scale AI generated omnidirectional image IQA database named AIGCOIQA2024 and constructing a comprehensive benchmark. We first generate 300 omnidirectional images based on 5 AIGC models utilizing 25 text prompts. A subjective IQA experiment is conducted subsequently to assess human visual preferences from three perspectives including quality, comfortability, and correspondence. Finally, we conduct a benchmark experiment to evaluate the performance of state-of-the-art IQA models on our database. The AIGCOIQA2024 database is released to facilitate future research on https://github.com/IntMeGroup/AIGCOIQA.
Huiyu Duan, Yucheng Zhu, Xiaohong Liu 0001, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
ICIP9
2024 "Discriminability-Experimental Cost" Tradeoff in Subjective Video Quality Assessment of Codec: DCR with EVP Rating Scale Versus ACR-HR
abstract
This work uses naive observers to compare two subjective studies conducted in a controlled laboratory environment on SDR HD, UHD, and HDR UHD contents. These tests aim to compare the precision and accuracy of a modified Degradation Category Rating (DCR) and Absolute Category Rating with Hidden Reference (ACR-HR) subjective methods for video quality assessment. The modified version of the DCR method includes a repetition of both reference and distorted stimuli; and utilizes an 11-grade rating scale from Expert Viewing Protocol (EVP) of ITU-R BT.500-15 standards. In the second subjective protocol, ACR-HR operates without repetition and with the 5-grade quality scale from ITU standards. We extensively analyze the scale usage and compare Mean Opinion Score (MOS) discriminability in both subjective studies. We show that both methods can retrieve accurate MOS. However, the ACR-HR method achieves better discriminability among MOS than DCR with the EVP rating scale while reducing the experimental effort by a factor of two, i.e., the cost of the experiment. The findings of this work give new insight into how to perform cost-efficient subjective tests for video quality estimation with naive observers and how to retrieve good MOS estimates.
Andreas Pastor, Ioannis Katsavounidis, Lukas Krasula, Andrey Norkin, Hassene Tmar, Patrick Le Callet
PCS7
2024 Beyond Curves and Thresholds - Introducing Uncertainty Estimation to Satisfied User Ratios for Compressed Video
abstract
Just Noticeable Difference (JND) establishes the threshold between two images or videos wherein differences in quality remain imperceptible to an individual. This threshold, collectively known as the Satisfied User Ratio (SUR), holds significant importance in image and video compression applications, ensuring that differences in quality are imperceptible to the majority (p%) of users, known as p%SUR. While substantial efforts have been dedicated to predicting the p%SUR for various encoding parameters (e.g., QP) and quality metrics (e.g., VMAF), referred to as proxies, systematic consideration of the prediction uncertainties associated with these proxies has hitherto remained unexplored. In this paper, we analyze the uncertainty of p%SUR through Confidence Interval (CI) estimation and assess the consistency of various Video Quality Metrics (VQMs) as proxies for SUR. The analysis reveals challenges in directly using p%SUR as ground truth for training models and highlights the need for uncertainty estimation for SUR with different proxies.
Hadi Amirpour, Raimund Schatz, Patrick Le Callet, Christian Timmerer
PCS4
2024 Perceptual Skin Tone Color Difference Measurement for Portrait Photography
abstract
In portrait photography, measuring the perceptual color differences (CDs) of skin tone is significant. Many studies have documented that the perception of skin tone is characteristically different from that of other colors. However, most existing CD measures are proposed based on psychophysical data of uniform color patches or natural images, and do not generalize well to the measurement of skin tone. In this paper, we construct the first large-scale portrait dataset for perceptual skin tone CD assessment and conduct psychophysical experiments to collect 160,000 perceptual CD judgments for 40,000 image triplets. Based on this dataset, we propose a deep skin tone CD measure for portrait photography. Extensive experiments demonstrate that our measure substantially outperforms existing CD measures on the problem of assessing skin tone CDs. The constructed dataset and code will be released to facilitate future research.
Shiqi Gao, Huiyu Duan, Qihang Xu, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
VCIP7
2024 The Salient360! toolbox: Handling gaze data in 3D made easy
abstract
Eye tracking has historically been a very popular tool. The data it records allow us to understand how people behave and what they attend to within our visual world; under this perspective the experiments, applications and use-cases are endless. Therefore, it is not surprising to witness a strong rise in the use of eXtended Reality (XR) devices with embedded eye trackers in research. These devices allow for less obtrusive experimenting conditions, and a significantly higher experimental control compared to traditional desktop testing. The use of eye tracking in XR is increasing and so is the need for a toolbox enabling consensus about eye tracking methods in 3D. We present the Salient360! toolbox: it implements functions to identify saccades and fixations and output gaze features (e.g., saccade directions) to generate saliency maps, fixation maps, and scanpath data. It implements comparisons of gaze data with methods adapted to 3D. We plan continuous improvements of the toolbox as the community develops new tools and methods dedicated to 360°gaze tracking. We hope that this toolbox will spark discussions about the methodology of 3D gaze processing, facilitate running experiments, and improve studying gaze in 3D. https://github.com/David-Ef/salient360Toolbox
Erwan J. David, Jesús Gutiérrez 0001, Melissa Le-Hoa Vo, Antoine Coutrot, Matthieu Perreira Da Silva, Patrick Le Callet
Comput. Graph.6
2024 Deep Blind Image Quality Assessment Using Dynamic Neural Model With Dual-Order Statistics
abstract
Deep convolutional neural networks (CNNs) have increasingly become a prominent method for blind image quality assessment (BIQA). The process of quality assessment typically involves feature extraction, average-based pooling, and quality regression. Based on this process, as well as the consensus that the visual quality of an image mainly relies on its content and distortions, this work improves CNNs for BIQA in two ways. First, considering the content-awareness of visual quality perception, we incorporate content-awareness via a dynamic filtering module to extract content-adaptive features and a dynamic regression module to learn content-adaptive perception rules based on local content and global semantics. Second, considering distortion-sensitivity in visual quality perception, we introduce second-order global variance pooling and combine it with global average pooling (GAP). First-order pooling methods like GAP are limited in distinguishing complex distortions that cause local degradation while preserving global features. Thus, pooling with dual-order statistics enables a more distortion-sensitive and discriminative global representation. These two improvements result in a content-adaptive BIQA model with a dual-order global pooling mechanism, improving generalization on diverse images with varying contents and distortion types. Extensive experiments on synthetic and authentic distortion datasets demonstrate state-of-the-art performance of the proposed approach.
Zihan Zhou 0007, Jing Li 0026, Dexiang Zhong, Yong Xu 0007, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.5
2024 BASICS: Broad Quality Assessment of Static Point Clouds in a Compression Scenario
abstract
Point clouds have become increasingly prevalent in representing 3D scenes within virtual environments, alongside 3D meshes. Their ease of capture has facilitated a wide array of applications on mobile devices, from smartphones to autonomous vehicles. Notably, point cloud compression has reached an advanced stage and has been standardized. However, the availability of quality assessment datasets, which are essential for developing improved objective quality metrics, remains limited. In this paper, we introduce BASICS, a large-scale quality assessment dataset tailored for static point clouds. The BASICS dataset comprises 75 unique point clouds, each compressed with four different algorithms including a learning-based method, resulting in the evaluation of nearly 1500 point clouds by 3500 unique participants. Furthermore, we conduct a comprehensive analysis of the gathered data, benchmark existing point cloud quality assessment metrics and identify their limitations. By publicly releasing the BASICS dataset, we lay the foundation for addressing these limitations and fostering the development of more precise quality metrics.
Ali Ak, Emin Zerman, Maurice Quach, Aladine Chetouani, Aljoscha Smolic, Giuseppe Valenzise, Patrick Le Callet
IEEE Trans. Multim.7
2023 The Salient360! Toolbox: Processing, Visualising and Comparing Gaze Data in 3D
abstract
Eye tracking can serve as a gateway to studying the mind. For this reason it has been adopted by a diverse range of scientific communities. With the improvement of the quality of head-mounted virtual reality devices (HMDs) over the past 10 years, eye tracking has been added to capture gaze in immersive environments. The use of HMDs with eye tracking is increasing significantly and so is the need for a toolbox enabling consensus about eye tracking methods in 3D. We present the Salient360! toolbox: it implements functions to identify saccades and fixations and output gaze characteristics (e.g., fixation duration or saccade directions), to generate saliency maps, fixation maps, and scanpath data. It also implements routines made to compare gaze data that were adapted to 3D. We hope that this toolbox will spark discussions about the methodology of 3D gaze processing, facilitate running experiments, and improve the gaze study in 3D. https://github.com/David-Ef/salient360Toolbox
Erwan J. David, Jesús Gutiérrez 0001, Melissa Le-Hoa Vo, Antoine Coutrot, Matthieu Perreira Da Silva, Patrick Le Callet
ETRA6
2023 Could the BubbleView Metaphor be used to Infer Visual Attention on 3D Graphical Content?
abstract
Understanding the deployment of human gaze on 3D graphical objects is of critical importance in order to propose rich and complex 3D environments without strong latency nor rendering constraints. However, the data needed to study this gaze deployment can be costly and difficult to obtain, especially in the context of the Covid-19 pandemic where in-lab experiments are strongly discouraged. In order to alleviate these issues, we propose to use the BubbleView metaphor as a way of crowdsourcing visual attention data on 3D graphical content. In this paper, we question the adequacy of this method to provide a reliable proxy for visual attention in the context of 3D graphical objects. Moreover, we show how data obtained in this manner can be used to train visual saliency models, with only a slight tradeoff in performances compared to the use of ground-truth eye-tracking data.
Alexandre Bruckert, Mona Abid, Matthieu Perreira Da Silva, Patrick Le Callet
ICASSP4
2023 Estimating Uncertainty On Video Quality Metrics
abstract
Video Quality Metrics (VQM) are models used to predict the score that a user would give to the quality of a video visualization. They are widely used in video processing systems, for monitoring end-to-end quality or system troubleshooting for example. In these scenarios, the improvement is quantified based on a certain enhancement of a VQM score and trouble-detection is done based on a certain drop or threshold computed based on a VQM. Yet, whether such improvement or fault-detection is worth a significant increase in power consumption is questionable. Therefore, the goal of this work is to propose a method to predict the uncertainty of the quality metric. In this paper, we propose a framework to evaluate the confidence interval of a VQM for a given content using simple video features. We assess the performance of the framework by using the confidence intervals to predict if two videos are of similar or different quality and show that in most cases our approach performs better than just using a constant confidence interval.
Patrick Le Callet, Suiyi Ling, Haixiong Wang, Ioannis Katsavounidis, Zafar Shahid, Cosmin Stejerean
ICASSP2
2023 Recovering Quality Scores in Noisy Pairwise Subjective Experiments Using Negative Log-Likelihood
abstract
To gather larger datasets to train data-angry deep learning quality assessment models, crowdsourcing has become essential to recruit participants. These participants are asked their opinion by directly rating stimuli, e.g., using single or double stimulus methodologies, or indirectly by ranking stimuli or comparing distances as in the Maximum Likelihood Difference Scaling method. In crowdsourcing, participants’ behaviors and environmental distractions are not controlled. So, the researcher must pay attention to the answers’ reliability. Cleaning methods exist for direct annotation subjective methodologies. However, solutions for indirect annotation methods are limited. In this work, we propose a method based on the negative log-likelihood to detect spammers among participants from their answers. To demonstrate its use, we applied it in a quadruplet preference-based scenario. The proposed method requires low computation and can be integrated into active-sampling strategies, where annotations available per comparison are small. We demonstrate that our method is robust to various spammer behaviors and accurate by removing only spammers. It helps reduce the gap between data collected in in-lab conditions (i.e., no spammer) and through crowdsourcing: our method reduces estimated uncertainties around data-points by 50%, and RMSE between estimations from an in-lab experiment and the same experiment performed in crowdsourcing by 1.8.
Andreas Pastor, Lukas Krasula, Zhi Li 0001, Patrick Le Callet
ICIP5
2023 ZREC: Robust Recovery of Mean and Percentile Opinion Scores
abstract
Observer screening and subject opinion score recovery is essential for collecting a reliable QoE database. This paper proposes a new method, ZREC*, which uses Z-scores to estimate subject bias, inconsistency, and content ambiguity. Additionally, we propose Mean Opinion Score (MOS) recovery and Percentile Opinion Score (POS) recovery scheme based on the three estimated parameters. ZREC does not fully reject subjects, rather adjust their coefficients in the MOS/POS recovery, allowing for more efficient use of data collection. The estimated parameters of ZREC are highly correlated with more complex solver-based methods and standards. In addition, ZREC recovers MOS with smaller confidence intervals than the state of the art. Experimental results also demonstrate that using recovered pthPOS as ground truth during training improves the performance of Satisfied User Ratio (SUR) prediction.
Ali Ak, Patrick Le Callet, Sriram Sethuraman, Kumar Rahul
ICIP3
2023 Just Noticeable Difference-Aware Per-Scene Bitrate-Laddering for Adaptive Video Streaming
abstract
In video streaming applications, a fixed set of bitrate-resolution pairs (known as a bitrate ladder) is typically used during the entire streaming session. However, an optimized bitrate ladder per scene may result in (i) decreased storage or delivery costs or/and (ii) increased Quality of Experience. This paper introduces a Just Noticeable Difference (JND)-aware perscene bitrate ladder prediction scheme (JASLA) for adaptive video-on-demand streaming applications. JASLA predicts jointly optimized resolutions and corresponding constant rate factors (CRFs) using spatial and temporal complexity features for a given set of target bitrates for every scene, which yields an efficient constrained Variable Bitrate encoding. Moreover, bitrate-resolution pairs that yield distortion lower than one JND are eliminated. Experimental results show that, on average, JASLA yields bitrate savings of 34.42% and 42.67% to maintain the same PSNR and VMAF, respectively, compared to the reference HTTP Live Streaming (HLS) bitrate ladder Constant Bitrate encoding using x265 HEVC encoder, where the maximum resolution of streaming is Full HD (1080p). Moreover, a 54.34% average cumulative decrease in storage space is observed.
Vignesh V. Menon, Prajit T. Rajendran, Hadi Amirpour, Patrick Le Callet, Christian Timmerer
ICME5
2023 Towards Guidelines for Subjective Haptic Quality Assessment: A Case Study on Quality Assessment of Compressed Haptic Signals
abstract
Modern systems are multimodal (e.g., video, audio, smell), and haptic feedback provides the user with additional entertainment and sensory immersion. Standard recommendation groups extensively studied and focused on video and audio subjective quality assessment, especially in signal transmission. In that context, subjective quality assessment and Quality of Experience (QoE) of Haptic signals is at its infant age. We propose further analyzing the collected data from a recent subjective quality assessment campaign as part of the MPEG haptic standardization group. In particular, we are addressing the following questions: 1) How the emerging field of haptic signal QoE can benefit from existing efforts of video and audio quality assessment standards? 2) How to detect possible outliers or characterize the rater’s reliability? 3) How does the discriminability of haptic tests increases with the number of raters? Towards this goal, we question if traditional analysis as proposed for audio or video signal are suitable, as well as other state-of-the-art techniques. We also compare the discriminability of the haptics quality assessment tests with other modalities such as audio, video, and immersive content (360° contents). We propose recommendations on the number of raters required to meet the usual discriminability obtained for other perceptual modalities and how to process ratings to remove possible noise and biases. These results could feed future recommendations in standards such as BT500-14 or P.913 but for haptic signals.
Andreas Pastor, Patrick Le Callet
ICME2
2023 From Temporal-Evolving to Spatial-Fixing: A Keypoints-Based Learning Paradigm for Visual Robotic Manipulation
abstract
The current learning pipelines for robotics manipulation infer movement primitives sequentially along the temporal-evolving axis, which can result in an accumulation of prediction errors and subsequently cause the visual observations to fall out of the training distribution. This paper proposes a novel hierarchical behavior cloning approach which tries to dissociate standard behaviour cloning (BC) pipeline to two stages. The intuition of this approach is to eliminate accumu-lation errors using a fixed spatial representation. At first stage, a high-level planner will be employed to translate the initial observation of the scene into task-specific spatial waypoints. Then, a low-level robotic path planner takes over the task of guiding the robot by executing a set of pre-defined elementary movements or actions known as primitives, with the goal of reaching the previously predicted waypoints. Our hierarchical keypoints-based paradigm aims to simplify existing temporal-evolving approach to a more simple way: directly spatialize the whole sequential primitives as a set of 8D waypoints only from the very first observation. Plentiful experiments demon-strate that our paradigm can achieve comparable results with Reinforcement Learning (RL) and outperforms existing offline BC approaches, with only a single-shot inference from the initial observation. Code and models are available at: https://github.com/KevinRiou22/spatial-fixing-il
Kévin Riou, Kaiwen Dong, Kévin Subrin, Yanjing Sun, Patrick Le Callet
IROS5
2023 Bridge the Gap between Visual Difference Prediction Model and Just Noticeable Difference Subjective Datasets
abstract
In video compression applications, the term 75%SUR (Satisfied User Ratio) is used to describe the compression parameter with which only 75% of the users can not notice the difference between one compressed media and its source. 75%SUR is widely used in JND (Just Noticeable Difference) modeling as a common threshold to standardize differences in users' perceptions of JND location. Visible difference detection is an essential step in JND prediction. However, Visible Difference Predictors (VDP), as objective quality metrics, are usually calibrated and applied on media quality datasets, no study has yet trained or applied the VDP on JND datasets. In this work, we will explore the feasibility of using the VDP model in predicting SUR and JND. We focus on Video Wise JND(VW-JND) and propose the model Extend-FvVDP, which maps the continuous quality scores output from the current best-performing VDP model, the FovVideoVDP, to VW-JND ground truth. Finally, Extend-FvVDP got a mean SUR prediction error of 0.0624, a mean JND prediction error of 1.9318. Our results show that VDP still performs on the JND datasets, and the JND prediction using VDP has the potential to exceed that of pure deep learning models.
Patrick Le Callet
MMSP3
2023 Perceptual annotation of local distortions in videos: tools and datasets
abstract
To assess the quality of multimedia content, create datasets, and train objective quality metrics, one needs to collect subjective opinions from annotators. Different subjective methodologies exist, from direct rating with single or double stimuli to indirect rating with pairwise comparisons. Triplet and quadruplet-based comparisons are a type of indirect rating. From these comparisons and preferences on stimuli, we can place the assessed stimuli on a perceptual scale (e.g., from low to high quality). Maximum Likelihood Difference Scaling (MLDS) solver is one of these algorithms working with triplets and quadruplets. A participant is asked to compare intervals inside pairs of stimuli: (a,b) and (c,d), where a,b,c,d are stimuli forming a quadruplet. However, one limitation is that the perceptual scales retrieved from stimuli of different contents are usually not comparable. We previously offered a solution to measure the inter-content scale of multiple contents. This paper presents an open-source python implementation of the method and demonstrates its use on three datasets collected in an in-lab environment. We compared the accuracy and effectiveness of the method using pairwise, triplet, and quadruplet for intra-content annotations. The code is available here: https://github.com/andreaspastor/MLDS_inter_content_scaling.
Andreas Pastor, Patrick Le Callet
MMSys2
2023 LLM-Based Interaction for Content Generation: A Case Study on the Perception of Employees in an IT Department
abstract
In the past years, AI has seen many advances in the field of NLP. This has led to the emergence of LLMs, such as the now famous GPT-3.5, which revolutionise the way humans can access or generate content. Current studies on LLM-based generative tools are mainly interested in the performance of such tools in generating relevant content (code, text or image). However, ethical concerns related to the design and use of generative tools seem to be growing, impacting the public acceptability for specific tasks. This paper presents a questionnaire survey to identify the intention to use generative tools by employees of an IT company in the context of their work. This survey is based on empirical models measuring intention to use (TAM by Davis, 1989, and UTAUT2 by Venkatesh and al., 2008). Our results indicate a rather average acceptability of generative tools, although the more useful the tool is perceived to be, the higher the intention to use seems to be. Furthermore, our analyses suggest that the frequency of use of generative tools is likely to be a key factor in understanding how employees perceive these tools in the context of their work. Following on from this work, we plan to investigate the nature of the requests that may be made to these tools by specific audiences.
Alexandre Agossah, Frédérique Krupa, Matthieu Perreira Da Silva, Patrick Le Callet
IMX4
2023 Video Consumption in Context: Influence of Data Plan Consumption on QoE
abstract
User expectations are one of the main factors on providing satisfactory QoE for streaming service providers. Measuring acceptability and annoyance of video content, therefore, provide a valuable insight when measured under a given context. In this ongoing work, we measure video QoE in terms of acceptability and annoyance for the remaining data in a mobile data plan context.. We show that simple logos can be used during the experiment to prompt the context to subjects and the different context levels may impact the user expectations and consequently their satisfactions. Finally, we show that objective metrics can be used to determine the acceptability and annoyance thresholds for a given context.
Ali Ak, Anne-Flore Perrin, Denise Noyes, Ioannis Katsavounidis, Patrick Le Callet
IMX5
2023 A Dataset of Gaze and Mouse Patterns in the Context of Facial Expression Recognition
abstract
Facial expression recognition is an important and challenging task for both the computer vision and affective computing communities, and even more specifically in the context of multimedia applications, where audience understanding is of particular interest. Recent data-oriented approaches have created the need for large-scale annotated datasets. However, most existing datasets present some weaknesses, because of the collecting methods used. In order to further highlight these issues, we investigate in this work how human visual attention is deployed when performing a facial expression recognition task. To do so, we carried out several complementary experiments, using the eye-tracking technology, as well as the BubbleView metaphor, both under laboratory and crowdsourcing settings. We show significant variations in gaze patterns depending on the emotion represented, but also on the difficulty of the task, i.e., whether the emotion is correctly recognised or not. Moreover, we use these results to propose recommendations on the ways to collect label data for facial expression recognition datasets.
Alexandre Bruckert, Lucie Lévêque, Matthieu Perreira Da Silva, Patrick Le Callet
IMX4
2023 Kinetic particles : from human pose estimation to an immersive and interactive piece of art questionning thought-movement relationships
abstract
Digital tools offer extensive solutions to explore novel interactive-art paradigms, by relying on various sensors to create installations and performances where the human activity can be captured, analysed and used to generate visual and sound universes in real-time. Deep learning approaches, including human detection and human pose estimation, constitute ideal human-art interaction mediums, as they allow automatic human gesture analysis, which can be directly used to produce the interactive piece of art. In this context, this paper presents an interactive work of art that explores the relationship between thought and movement by combining dance, philosophy, numerical arts, and deep learning. We present a novel system that combines a multi-camera setup to capture human movement, state-of-the-art human pose estimation models to automatically analyze this movement, and an immersive 180° projection system that projects a dynamic textual content that intuitively responds to the users’ behaviors. The demonstration being proposed consists of two parts. Firstly, a professional dancer will utilize the proposed setup to deliver a conference-show. Secondly, the audience will be given the opportunity to experiment and discover the potential of the proposed setup, which has been transformed into an interactive installation. This allows multiple spectators to engage simultaneously with clusters of words and letters extracted from the conference text.
Mickael Lafontaine, Julie Cloarec-Michaud, Kévin Riou, Kaiwen Dong, Patrick Le Callet
IMX6
2023 Use of immersive and interactive systems to objectively and subjectively characterize user experience in work places
abstract
This demo’s objective is to display how simulated environments can be used in the evaluation of work-related indoor environments compared to physical environments. In fact, in indoor environment evaluations, Virtual Reality (VR) offers new possibilities for experimental design as well as a functional rapprochement between the laboratory and real life. With VR, environmental parameters (e.g light, color, furniture...) can be easily manipulated at reasonable costs, allowing to control and guide the user’s sensorial experience. One main challenge is to acknowledge to which extent simulated environments are ecologically valid and which functions would be more solicited in different environmental-simulation display formats. User-centric evaluation and sensory analysis in the building sector is in its beginning; this new tool could be of benefit for the building sector, on one hand for methodological facilitation purposes and on the other for cost reductions. In order to achieve the objectives of this project, a first step is to develop and validate the indoor simulations. In environmental simulations, one of the most used formats, for its visual realism and ease of use are 360° panoramic photos and videos, which permits capturing physical-world images. In an objective of validation of the format, 360° photos of workplaces were taken in the building of Halle 6 Ouest of Nantes University and an immersive and interactive test based on physiological indicators to subjectively and objectively assess comfort and performances in work offices was developed. The demo will comprise a head-mounted display with integrated eye-tracking and the measure of electrodermal activity, heart rate and galvanic skin response.
Charlotte Scarpa, Gwénaëlle Haese, Toinon Vigier, Patrick Le Callet
IMX4
2023 Construction of immersive and interactive methodology based on physiological indicators to subjectively and objectively assess comfort and performances in work offices
abstract
The building sector and the indoor environment conception is undergoing major changes. There is a need to reconsider the way offices are built from a user’s centric point of view. Research has shown the influence of perceived comfort and satisfaction on performance in the workplace. By understanding how multi-sensory information is integrated into the nervous system and which environmental parameters influence the most perception, it could be possible to improve work environments. With the emergence of new virtual reality (VR) and augmented reality (AR) technologies, the collection and processing of sensory information is rapidly advancing, moving forward more dynamic aspects of sensory perception. Through simulated environments, environmental parameters can be easily manipulated at reasonable costs, allowing control and guiding the user’s sensory experience. Moreover, the effects of contextual and surrounding stimuli on users can be easily collected throughout the test, in the form of physiological and behavioral data. Through the use of indoor simulations, this doctoral research goal is to develop a multi-criteria comfort scale based on physiological indicators under performance constraints. In doing this, it would be possible to define new quality indicators combining the different physical factors adapted to the uses and space. In order to achieve the objectives of this project, the first step is to develop and validate an immersive and interactive methodology for the assessment of multisensory information on comfort and performance in work environments.
Charlotte Scarpa, Gwénaëlle Haese, Toinon Vigier, Patrick Le Callet
IMX4
2023 Subjective Test Environments: A Multifaceted Examination of Their Impact on Test Results
abstract
Quality of Experience (QoE) in video streaming scenarios is significantly affected by the viewing environment and display device. Understanding and measuring the impact of these settings on QoE can help develop viewing environment-aware metrics and improve the efficiency of video streaming services. In this ongoing work, we conducted a subjective study in both laboratory and home settings using the same content and design to measure QoE in Degradation Category Rating (DCR). We first analyzed subject inconsistency and confidence intervals of the Mean Opinion Scores (MOS) between the two settings. We then used statistical models such as ANOVA and t-test to analyze the differences in subjective tests on video quality between the two viewing environments. Additionally, we employed the Eliminated-By-Aspects (EBA) model to quantify the influence of different settings on the measured QoE. We conclude with several research questions that could be further explored to better understand the impact of the viewing environment on QoE.
Ali Ak, Charles Dormeval, Patrick Le Callet, Kumar Rahul, Sriram Sethuraman
IMX4
2023 Predicting local distortions introduced by AV1 using Deep Features
abstract
Semantics extracted by filters in deep learning networks correlate well with how human eyes perceive distortions. These methods (e.g., LPIPS, PieAPP, etc.) rely on the relative difference in activation between feature maps in pairs of references and distorted patches. However, Deep Feature extraction can be expensive to compute as a difference of latent code between reference and distorted frames. Therefore, it is challenging to integrate them into the decision process of modern video codecs like AV1, making thousands of encoding trials during exhaustive Rate-Distortion Optimization (RDO) searches. In this study, we present a method using deep features to predict the distortion perceived locally by human eyes in AV1-encoded videos. The prediction relies on Deep Features extracted from the reference frame only to weigh the Mean Squared Error (MSE) introduced during encoding. This approach will make integration into video codecs easier as a pre-processing step before starting encoding. We show the superiority of the proposed metric against other Reference-Only metrics on a dataset of local distortions in videos. We achieve comparable performance as state-of-the-art Full-Reference video quality metrics.
Andreas Pastor, Lukas Krasula, Zhi Li 0001, Patrick Le Callet
VCIP5
2023 Rethinking Scene Graphs for Action Recognition
abstract
Over the last years, Graph Neural Networks (GNNs) have been widely used in a variety of applications, including action recognition. Scene graphs are extracted from videos and fed to a GNN in order to predict the action represented. However, in previous works, choices regarding the design of such scene graphs are often arbitrary; for instance, directed temporal edges are added without giving the GNN the capacity to use this information. In this work, we rethink the way scene graphs are built, taking inspiration from line graphs in order to propose a new design that can be applied to any type of human activity. We perform our experiments on 2 datasets and show that adapting our GNN so that it can make use of temporal edges improves its precision up to 7.5% for action recognition. We also show that adopting our alternate design for scene graphs further improves performance by an additional 14%, bringing new perspectives to this field.
Mathieu Riand, Patrick Le Callet, Laurent Dollé
VCIP2
2023 Enhancing Satisfied User Ratio (SUR) Prediction for VMAF Proxy through Video Quality Metrics
abstract
In adaptive video streaming, optimizing the selection of representations for the encoding bitrate ladder has a significant impact on the quality and economics of media delivery. An efficient way to select representations for the bitrate ladder of a given clip is to consider the Satisfied User Ratio (SUR) of the perceived quality of consecutive representations. This ensures that only representations with one Just Noticeable Difference (JND) are encoded and streamed by avoiding encoding similar-quality representations. VMAF (Video Multi-method Assessment Fusion) presently stands as the most commonly utilized quality metric for constructing bitrate ladders. Hence, the precise determination of JND-optimal encoding step-sizes for the VMAF proxy holds paramount importance; nevertheless, this task is intricate and can present considerable challenges.In this paper, we evaluate the effectiveness of different Video Quality Metrics (VQMs) in predicting SUR for the VMAF proxy to better capture content-specific characteristics. Our experimental results provide evidence that incorporating VQMs can improve the precision of the SUR prediction for the VMAF proxy. Compared to a state-of-the-art approach that utilizes video complexity metrics, our proposed approach, which incorporates two quality metrics—specifically, VMAF and SSIM calculated at an optimized quantization parameter (QP)—achieves a substantially reduced Mean Absolute Error (MAE) of 1.67. In contrast, the state-of-the-art approach yields an MAE of 2.01. Hence, we recommend using the above quality metrics to improve the accuracy of the SUR prediction for the VMAF proxy.
Hadi Amirpour, Raimund Schatz, Christian Timmerer, Patrick Le Callet
VCIP5
2023 RV-TMO: Large-Scale Dataset for Subjective Quality Assessment of Tone Mapped Images
abstract
Tone mapping operators (TMO) are functions that map high dynamic range (HDR) images to a standard dynamic range (SDR), while aiming to preserve the perceptual cues of a scene that govern its visual quality. Despite the increasing number of studies on quality assessment of tone mapped images, current subjective quality datasets have relatively small numbers of images and subjective opinions. Moreover, existing challenges in transferring laboratory experiments to crowdsourcing platforms put a barrier for collecting large-scale datasets through crowdsourcing. In this work, we address these challenges and propose the RealVision-TMO (RV-TMO), a large-scale tone mapped image quality dataset. RV-TMO contains 250 unique HDR images, their tone mapped versions obtained using four TMOs and pairwise comparison results from seventy unique observers for each pair. To the best of our knowledge, this is the largest dataset available in the literature for quality evaluation of TMOs by the number of tone mapped images and number of annotations. Furthermore, we provide a content selection strategy to identify interesting and challenging HDR images. We also propose a novel methodology for observer screening in pairwise experiments. Our work does not only provide annotated data to benchmark existing objective quality metrics, but also paves the path to building new metrics for tone mapping quality evaluation.
Ali Ak, Abhishek Goswami, Wolf Hauser, Patrick Le Callet, Frédéric Dufaux
IEEE Trans. Multim.4
2023 Textured Mesh Quality Assessment: Large-scale Dataset and Deep Learning-based Quality Metric
abstract
Over the past decade, three-dimensional (3D) graphics have become highly detailed to mimic the real world, exploding their size and complexity. Certain applications and device constraints necessitate their simplification and/or lossy compression, which can degrade their visual quality. Thus, to ensure the best Quality of Experience, it is important to evaluate the visual quality to accurately drive the compression and find the right compromise between visual quality and data size. In this work, we focus on subjective and objective quality assessment of textured 3D meshes. We first establish a large-scale dataset, which includes 55 source models quantitatively characterized in terms of geometric, color, and semantic complexity, and corrupted by combinations of five types of compression-based distortions applied on the geometry, texture mapping, and texture image of the meshes. This dataset contains over 343k distorted stimuli. We propose an approach to select a challenging subset of 3,000 stimuli for which we collected 148,929 quality judgments from over 4,500 participants in a large-scale crowdsourced subjective experiment. Leveraging our subject-rated dataset, a learning-based quality metric for 3D graphics was proposed. Our metric demonstrates state-of-the-art results on our dataset of textured meshes and on a dataset of distorted meshes with vertex colors. Finally, we present an application of our metric and dataset to explore the influence of distortion interactions and content characteristics on the perceived quality of compressed textured meshes.
Yana Nehmé, Johanna Delanoy, Florent Dupont, Jean-Philippe Farrugia, Patrick Le Callet, Guillaume Lavoué
ACM Trans. Graph.5
2022 Considering User Agreement in Learning to Predict the Aesthetic Quality
abstract
How to robustly rank the aesthetic quality of given images has been a long-standing ill-posed topic. Such challenge stems mainly from the diverse subjective opinions of different observers about the varied types of content. There is a growing interest in estimating the user agreement by considering the standard deviation (σ) of the scores, instead of only predicting the mean aesthetic opinion score (µ). Nevertheless, when comparing a pair of contents, few studies consider how confident are we regarding the difference in the aesthetic scores. In this paper, we thus propose (1) a re-adapted multi-task attention network to predict both the mean opinion score and the standard deviation in an end-to-end manner; (2) a brand-new confidence interval ranking loss that encourages the model to focus on image-pairs that are less certain about the difference of their aesthetic scores. With such loss, the model is encouraged to learn the uncertainty of the content that is relevant to the diversity of observers’ opinions, i.e., user disagreement. Extensive experiments have demonstrated that the proposed multi-task aesthetic model achieves state-of-the-art performance on two different types of aesthetic datasets, i.e., AVA and TMGA.
Suiyi Ling, Andreas Pastor, Junle Wang, Patrick Le Callet
ICASSP4
2022 Improving Maximum Likelihood Difference Scaling Method To Measure Inter Content Scale
abstract
The goal of most subjective studies is to place a set of stimuli on a perceptual scale. This is mostly done directly by rating, e.g. using single or double stimulus methodologies, or indirectly by ranking or pairwise comparison. All these methods estimate the perceptual magnitudes of the stimuli on a scale. However, procedures such as Maximum Likelihood Difference Scaling (MLDS) have shown that considering perceptual distances can bring benefits in terms of discriminatory power, observers’ cognitive load, and the number of trials required. One of the disadvantages of the MLDS method is that the perceptual scales obtained for stimuli created from different source content are generally not comparable. In this paper, we propose an extension of the MLDS method that ensures inter-content comparability of the results and shows its usefulness especially in the presence of observer errors.
Andreas Pastor, Lukas Krasula, Zhi Li 0001, Patrick Le Callet
ICASSP5
2022 Specialised Video Quality Model For Enhanced User Generated Content (UGC) With Special Effects
abstract
User Generated Content (UGC) refers to media generated by users for end-consumers that represent most of the media exchange on social media. UGC is subject to acquisition and transmission limitations that disable access to the pristine, i.e., perfect source content. Evaluating their quality, especially with current pre- and post-processing algorithms or filters, is a major issue for most off-the-shelf full-reference quality metrics. We propose to conduct a benchmark on existing full-reference, non-reference, and aesthetic quality metrics for UGC with special effects. We aim to identify the challenges posed by both UGC and filtering. We then propose a new combination of metrics tailored to enhanced and filtered UGC, which reaches a trade-off between complexity and accuracy.
Anne-Flore Perrin, Yejing Xie, Yiting Liao, Patrick Le Callet
ICASSP6
2022 Subjective And Objective Quality Assessment Of Mobile Gaming Video
abstract
Nowadays, with the vigorous expansion and development of gaming video streaming techniques and services, the expectation of users, especially the mobile phone users, for higher quality of experience is also growing swiftly. As most of the existing research focuses on traditional video streaming, there is a clear lack of both subjective study and objective quality models that are tailored for quality assessment of mobile gaming content. To this end, in this study, we first present a brand new Tencent Gaming Video dataset containing 1293 mobile gaming sequences encoded with three different codecs. Second, we propose an objective quality framework, namely Efficient hard-RAnk Quality Estimator (ERAQUE), that is equipped with (1) a novel hard pairwise ranking loss, which forces the model to put more emphasis on differentiating similar pairs; (2) an adapted model distillation strategy, which could be utilized to compress the proposed model efficiently without causing significant performance drop. Extensive experiments demonstrate the efficiency and robustness of our model.
Shaoguo Wen, Suiyi Ling, Junle Wang, Yanqing Jing, Patrick Le Callet
ICASSP6
2022 On the Accuracy of Open Video Quality Metrics for Local Decision in AV1 Video Codec
abstract
VMAF is a popular objective quality metric used for video quality evaluation. The power of VMAF has been demonstrated for a wide variety of video scales and encoding processes. However, its ability to evaluate the quality of small video patches has not yet been tested, despite its importance for encoding algorithms. We applied Maximum Likelihood Difference Scaling (MLDS) methodology to estimate supra-threshold perceptual differences in localized sections in videos, also known as tubes, encoded using AV1. We further used the results to assess the performance of VMAF in this scenario and proposed a recalibration of the algorithm to improve its agreement with the subjective data.
Andreas Pastor, Lukas Krasula, Zhi Li 0001, Patrick Le Callet
ICIP5
2022 When is the Cleaning of Subjective Data Relevant to Train UGC Video Quality Metrics?
abstract
Outlier analysis and spammer detection recently gained momentum in order to reduce uncertainty of subjective ratings in image & video quality assessment tasks. The large proportion of unreliable ratings from online crowdsourcing experiments and the need for qualitative and quantitative large-scale studies in the deep-learning ecosystem played a role in this event. We study the effect that data cleaning has on trainable models predicting the visual quality for videos, and present results demonstrating when cleaning is necessary to reach higher efficiency. To this end, we present and analyze a benchmark on clean and noisy User Generated Content (UGC) large-scale datasets on which we re-trained models, followed by an empirical exploration of the constraint of data removal. Our results show that a dataset presenting between 7 and 30% of outliers benefits from cleaning before training.
Anne-Flore Perrin, Charles Dormeval, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Patrick Le Callet
ICIP6
2022 On The Benefit of Parameter-Driven Approaches for the Modeling and the Prediction of Satisfied User Ratio for Compressed Video
abstract
The human eye cannot perceive small pixel changes in images or videos until a certain threshold of distortion. In the context of video compression, Just Noticeable Difference (JND) is the smallest distortion level from which the human eye can perceive the difference between reference video and the distorted/compressed one. Satisfied-User-Ratio (SUR) curve is the complementary cumulative distribution function of the individual JNDs of a viewer group. However, most of the previous works predict each point in SUR curve by using features both from source video and from compressed videos with assumption that the group-based JND annotations follow Gaussian distribution, which is neither practical nor accurate. In this work, we firstly compared various common functions for SUR curve modeling. Afterwards, we proposed a novel parameter-driven method to predict the video-wise SUR from video features. Besides, we compared the prediction results of source-only features based (SRC-based) models and source plus compressed videos features (SRC+PVS-based) models.
Patrick Le Callet, Anne-Flore Perrin, Sriram Sethuraman, Kumar Rahul
ICIP2
2022 Reinforcement Learning Based Point-Cloud Acquisition and Recognition Using Exploration-Classification Reward Combination
abstract
3D points acquisitions based on robust sensors such as tactile or laser sensors are true alternatives to computer vision for 3D object recognition. In real life scenarios where robots are equipped with such sensors to acquire 3D data, only few points can be iteratively collected in a reasonable amount of time. However, existing Point-cloud classifiers are extremely sensitive to sparse points, missing parts and noise. To compensate for the sparsity of the data, some Reinforcement Learning (RL) based approaches have been proposed to learn a sparse yet efficient exploration of the target object regarding the 3D recognition objective. However, existing RL approaches only focus on classification performances to guide the training of the active acquisition-and-classification frameworks, and thus fail to dissociate poor exploration strategy (missing parts, noisy points) from actual classifier mistakes on proper data. In this study, we proposed a new RL framework that was rewarded regarding both the classification performances and the exploration quality. Our trained framework outperforms existing State-Of-The-Art models on 3D geometric objects classification. We further showed that our trained framework learnt to alternate between (1) a clean and broad exploration strategy, suitable for easily distinguishable categories, and (2) a specific local exploration strategy, facilitating the discrimination of similar categories.
Kévin Riou, Kévin Subrin, Patrick Le Callet
ICME3
2022 Imitation from Observation using RL and Graph-based Representation of Demonstrations
abstract
Teaching robots behavioral skills by leveraging examples provided by an expert, also referred to as Imitation Learning from Observation (IfO or ILO), is a promising approach for learning novel tasks without requiring a task-specific reward function to be engineered. We propose a RL-based framework to teach robots manipulation tasks given expert observation-only demonstrations. First, a representation model is trained to extract spatial and temporal features from demonstrations. Graph Neural Networks (GNNs) are used to encode spatial patterns, while LSTMs and Transformers are used to encode temporal features. Second, based on an off-the-shelf RL algorithm, the demonstrations are leveraged through the trained representation to guide the policy training towards solving the task demonstrated by the expert. We show that our approach compares favorably to state-of-the-art IfO algorithms with a 99% success rate and transfers well to the real world.
Yassine El Manyari, Patrick Le Callet, Laurent Dollé
ICMLA2
2022 QoEVMA'22: 2nd Workshop on Quality of Experience (QoE) in Visual Multimedia Applications
abstract
Nowadays, people spend dramatically more time on watching videos through different devices. The advanced hardware technology and network allow for the increasing demands of users viewing experience. Thus, enhancing the Quality of Experience of end-users in advanced multimedia is the ultimate goal of service providers, as good services would attract more consumers. Quality assessment is thus important. The second workshop on "Quality of Experience (QoE) in visual multimedia applications" (QoEVMA'22) focuses on the QoE assessment of any visual multimedia applications both subjectively and objectively. The topics include 1) QoE assessment on different visual multimedia applications, including VoD for movies, dramas, variety shows, UGC on social networks, live streaming videos for gaming/shopping/social, etc. 2) QoE assessment for different video formats in multimedia services, including 2D, stereoscopic 3D, High Dynamic Range (HDR), Augmented Reality (AR), Virtual Reality (VR), 360, Free-Viewpoint Video(FVV), etc. 3) Key performance indicators (KPI) analysis for QoE. This summary gives a brief overview of the workshop, which took place on October 14, 2022 in Lisbon, Portugal, as a half-day workshop. The complete QOEVMA'22 workshop proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3552469
Jing Li 0026, Patrick Le Callet, Xinbo Gao 0001, Zhi Li 0001, Wen Lu 0004, Junle Wang
ACM Multimedia2
2022 Perception of video quality at a local spatio-temporal horizon: research proposal
abstract
This paper contains the research proposal of Andréas Pastor that was presented at the MMSys 2022 doctoral symposium. Encoding video for streaming on Internet has become a major topic to reduce the consumption of bandwidth and latency. At the same time, the human perception of distortions has been explored in multiple research projects, especially for distortions generated by Coder-DECoder (CODEC) algorithms. These algorithms operate in a rate-distortion optimization paradigm to efficiently compress video content. This optimization can be driven by metrics that are most of the time not based on the human perception, and more importantly, not tuned to reflect the local perception of distortions by human eyes.
Andreas Pastor, Patrick Le Callet
MMSys2
2022 Just noticeable difference (JND) and satisfied user ratio (SUR) prediction for compressed video: research proposal
abstract
This paper contains the research proposal of Jingwen ZHU that was presented at the MMSys 2022 doctoral symposium. Just noticeable difference (JND) is the minimum amount of distortion from which human eyes can perceive difference between the original stimuli and distorted stimuli. With the rapid raise of multimedia demand, it is crucial to apply JND into the visual communication systems to use the least resources (e.g., bandwidth and storage) but without damaging the Quality of Experience (QoE) of end-users. In this thesis, we focus on the JND prediction for compressed video to guide the choice of optimal encoding parameters for video streaming service. In this paper, we analyse the limitations of the current JND prediction models and present five main research questions to address these challenges.
Patrick Le Callet
MMSys2
2022 Subjective test methodology optimization and prediction framework for Just Noticeable Difference and Satisfied User Ratio for compressed HD video
abstract
Just Noticeable Difference (JND) and Satisfied User Ratio (SUR) has been widely investigated for compressed image and video to use the least resources (e.g., storage and bandwidth) without damaging the Quality of Experience (QoE) for end users. However, the current JND subjective test methodologies are extremely time consuming due to the large range of encoding parameters. Besides, the state-of-the-arts SUR/JND prediction models get non-negligible prediction error due to the limited masking effect features. To this end, we first proposed a preprocessing method to reduce the JND subjective test time by using dynamic range of encoding parameters and collected a new Video-Wise JND (VW-JND) datasets for HD videos: HDVJND. Afterwards, based on the collected datasets, we proposed a SUR prediction framework by extracting 3 types of features 1) masking effect features; 2) bitstreams features; 3) content features. Feature selection is applied to extracted features before regression. Besides, we also compared the direct and indirect SUR value predictions methods. Experiment results shows that our proposed optimization can reduce 7.14% of the subjective experiment time compared to the widely used Robust Binary Search (RBS). Furthermore, the proposed SUR and JND prediction frameworks outperform the SOTA model in HD-VJND datasets.
Anne-Flore Perrin, Patrick Le Callet
PCS3
2022 Confusing Image Quality Assessment: Toward Better Augmented Reality Experience
abstract
With the development of multimedia technology, Augmented Reality (AR) has become a promising next-generation mobile platform. The primary value of AR is to promote the fusion of digital contents and real-world environments, however, studies on how this fusion will influence the Quality of Experience (QoE) of these two components are lacking. To achieve better QoE of AR, whose two layers are influenced by each other, it is important to evaluate its perceptual quality first. In this paper, we consider AR technology as the superimposition of virtual scenes and real scenes, and introduce visual confusion as its basic theory. A more general problem is first proposed, which is evaluating the perceptual quality of superimposed images, i.e., confusing image quality assessment. A ConFusing Image Quality Assessment (CFIQA) database is established, which includes 600 reference images and 300 distorted images generated by mixing reference images in pairs. Then a subjective quality perception experiment is conducted towards attaining a better understanding of how humans perceive the confusing images. Based on the CFIQA database, several benchmark models and a specifically designed CFIQA model are proposed for solving this problem. Experimental results show that the proposed CFIQA model achieves state-of-the-art performance compared to other benchmark models. Moreover, an extended ARIQA study is further conducted based on the CFIQA study. We establish an ARIQA database to better simulate the real AR application scenarios, which contains 20 AR reference images, 20 background (BG) reference images, and 560 distorted images generated from AR and BG references, as well as the correspondingly collected subjective quality ratings. Three types of full-reference (FR) IQA benchmark variants are designed to study whether we should consider the visual confusion when designing corresponding IQA algorithms. An ARIQA metric is finally proposed for better evaluating the perceptual quality of AR images. Experimental results demonstrate the good generalization ability of the CFIQA model and the state-of-the-art performance of the ARIQA model. The databases, benchmark models, and proposed metrics are available at: https://github.com/DuanHuiyu/ARIQA.
Huiyu Duan, Xiongkuo Min, Yucheng Zhu, Guangtao Zhai, Xiaokang Yang 0001, Patrick Le Callet
IEEE Trans. Image Process.6
2022 Quality Assessment of DIBR-Synthesized Views Based on Sparsity of Difference of Closings and Difference of Gaussians
abstract
Images synthesized using depth-image-based-rendering (DIBR) techniques may suffer from complex structural distortions. The goal of the primary visual cortex and other parts of brain is to reduce redundancies of input visual signal in order to discover the intrinsic image structure, and thus create sparse image representation. Human visual system (HVS) treats images on several scales and several levels of resolution when perceiving the visual scene. With an attempt to emulate the properties of HVS, we have designed the no-reference model for the quality assessment of DIBR-synthesized views. To extract a higher-order structure of high curvature which corresponds to distortion of shapes to which the HVS is highly sensitive, we define a morphological oriented Difference of Closings (DoC) operator and use it at multiple scales and resolutions. DoC operator nonlinearly removes redundancies and extracts fine grained details, texture of an image local structure and contrast to which HVS is highly sensitive. We introduce a new feature based on sparsity of DoC band. To extract perceptually important low-order structural information (edges), we use the non-oriented Difference of Gaussians (DoG) operator at different scales and resolutions. Measure of sparsity is calculated for DoG bands to get scalar features. To model the relationship between the extracted features and subjective scores, the general regression neural network (GRNN) is used. Quality predictions by the proposed DoC-DoG-GRNN model show higher compatibility with perceptual quality scores in comparison to the tested state-of-the-art metrics when evaluated on four benchmark datasets with synthesized views, IRCCyN/IVC image/video dataset, MCL-3D stereoscopic image dataset and IST image dataset.
Dragana Sandic-Stankovic, Dragan Kukolj, Patrick Le Callet
IEEE Trans. Image Process.3
2022 UIF: An Objective Quality Assessment for Underwater Image Enhancement
abstract
Due to complex and volatile lighting environment, underwater imaging can be readily impaired by light scattering, warping, and noises. To improve the visual quality, Underwater Image Enhancement (UIE) techniques have been widely studied. Recent efforts have also been contributed to evaluate and compare the UIE performances with subjective and objective methods. However, the subjective evaluation is time-consuming and uneconomic for all images, while existing objective methods have limited capabilities for the newly-developed UIE approaches based on deep learning. To fill this gap, we propose an Underwater Image Fidelity (UIF) metric for objective evaluation of enhanced underwater images. By exploiting the statistical features of these images in CIELab space, we present the naturalness, sharpness, and structure indexes. Among them, the naturalness and sharpness indexes represent the visual improvements of enhanced images; the structure index indicates the structural similarity between the underwater images before and after UIE. We combine all indexes with a saliency-based spatial pooling and thus obtain the final UIF metric. To evaluate the proposed metric, we also establish a first-of-its-kind large-scale UIE database with subjective scores, namely Underwater Image Enhancement Database (UIED). Experimental results confirm that the proposed UIF metric outperforms a variety of underwater and general-purpose image quality metrics. The database and source code are available at https://github.com/z21110008/UIF.
Yannan Zheng, Rongfu Lin, Tiesong Zhao, Patrick Le Callet
IEEE Trans. Image Process.5
2022 SMGEA: A New Ensemble Adversarial Attack Powered by Long-Term Gradient Memories
abstract
Deep neural networks are vulnerable to adversarial attacks. More importantly, some adversarial examples crafted against an ensemble of source models transfer to other target models and, thus, pose a security threat to black-box applications (when attackers have no access to the target models). Current transfer-based ensemble attacks, however, only consider a limited number of source models to craft an adversarial example and, thus, obtain poor transferability. Besides, recent query-based black-box attacks, which require numerous queries to the target model, not only come under suspicion by the target model but also cause expensive query cost. In this article, we propose a novel transfer-based black-box attack, dubbed serial-minigroup-ensemble-attack (SMGEA). Concretely, SMGEA first divides a large number of pretrained white-box source models into several "minigroups." For each minigroup, we design three new ensemble strategies to improve the intragroup transferability. Moreover, we propose a new algorithm that recursively accumulates the "long-term" gradient memories of the previous minigroup to the subsequent minigroup. This way, the learned adversarial information can be preserved, and the intergroup transferability can be improved. Experiments indicate that SMGEA not only achieves state-of-the-art black-box attack ability over several data sets but also deceives two online black-box saliency prediction systems in real world, i.e., DeepGaze-II (https://deepgaze.bethgelab.org/) and SALICON (http://salicon.net/demo/). Finally, we contribute a new code repository to promote research on adversarial attack and defense over ubiquitous pixel-to-pixel computer vision tasks. We share our code together with the pretrained substitute model zoo at https://github.com/CZHQuality/AAA-Pix2pix.
Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Xiongkuo Min, Guodong Guo, Patrick Le Callet
IEEE Trans. Neural Networks Learn. Syst.8
2021 Combining Video Quality Metrics To Select Perceptually Accurate Resolution In A Wide Quality Range: A Case Study
abstract
It is beneficial to have adaptive resolution selection during encoding for the viewer’s optimal experience in a video delivery scenario. For instance, per-title content-adaptive techniques have been exploited for OTT/VOD delivery. An established solution is to build the selection of better resolution on typical video quality metrics. In this context, this paper first introduces measuring the performance of video quality metrics to accurately predict which encoding resolution is subjectively better suited to a particular scene of interest in a video. Then, with a Random Forest (RF) classifier, a novel, subjectively accurate classifier called RF-based fusion metric is proposed to decide which encoding resolution is best suited to compensate for individual metrics’ disparate performance over different quality ranges. The proposed RF classifier encompasses classical video quality metric scores as features and is trained to be closer to subjective experimental results. Performance comparison of this proposed RF-based fusion metric is discussed in comparison to other metrics and subjective experiments.
Madhukar Bhat, Jean-Marc Thiesse, Patrick Le Callet
ICIP3
2021 Cmdm-Vac: Improving A Perceptual Quality Metric For 3D Graphics By Integrating A Visual Attention Complexity Measure
abstract
Many objective quality metrics have been proposed over the years to automate the task of subjective quality assessment. However, few of them are designed for 3D graphical contents with appearance attributes; existing ones are based on geometry and color measures, yet they ignore the visual saliency of the objects. In this paper, we combined an optimal subset of geometry-based and color-based features, provided by a state-of-the-art quality metric for 3D colored meshes, with a visual attention complexity feature adapted to 3D graphics. The performance of our proposed new metric is evaluated on a dataset of 80 meshes with diffuse colors, generated from 5 source models corrupted by commonly used geometry and color distortions. With our proposed metric, we showed that the use of the attentional complexity feature brings a significant gain in performance and better stability.
Yana Nehmé, Mona Abid, Guillaume Lavoué, Matthieu Perreira Da Silva, Patrick Le Callet
ICIP5
2021 Seeing By Haptic Glance: Reinforcement Learning Based 3d Object Recognition
abstract
Human is able to conduct 3D recognition by a limited number of haptic contacts between the target object and his/her fingers without seeing the object. This capability is defined as ‘haptic glance’ in cognitive neuroscience. Most of the existing 3D recognition models were developed based on dense 3D data. Nonetheless, in many real-life use cases, where robots are used to collect 3D data by haptic exploration, only a limited number of 3D points could be collected. In this study, we thus focus on solving the intractable problem of how to obtain cognitively representative 3D key-points of a target object with limited interactions between the robot and the object. A novel reinforcement learning based framework is proposed, where the haptic exploration procedure (the agent iteratively predicts the next position for the robot to explore) is optimized simultaneously with the objective 3D recognition with actively collected 3D points. As the model is rewarded only when the 3D object is accurately recognized, it is driven to find the sparse yet efficient haptic-perceptua13D representation of the object. Experimental results show that our proposed model outperforms the state of the art models.
Kévin Riou, Suiyi Ling, Guillaume Gallot, Patrick Le Callet
ICIP4
2021 VVC partitioning decision driven by machine learning for a comprehensive hardware encoder
abstract
The new video coding standard VVC introduces Multi-Type Tree (MTT) partitioning, adding a horizontal and vertical splitting method. MTT significantly improves coding efficiency compared to QuadTree (QT) based partitioning used in HEVC, but with an exponential increase in the encoder’s computation complexity. Hence, MTT is bringing in a significant bottleneck for encoder design. The challenge is higher for realtime hardware encoders in silicon, bandwidth, power efficiency, and speed. This paper uses a CU-wise Convolutional Neural Network (CNN) based machine learning to drive the MTT partitioning decision. An application scenario of classifiers with minimal coding loss fitting a hardware encoder is presented. CNN’s are trained to predict possible partitioning at each depth of partitioning. The CNN networks are designed for real-time hardware scenario such that it has very light computational overhead to a hardware-based real-time encoder. The ML-driven partitioning system’s performance with various complexity presets are reported—the preset with the least BD-BR loss of 0.45% had a 33% reduction in computation complexity. The proposed low complexity preset reaches up to an 86% complexity reduction.
Madhukar Bhat, Jean-Marc Thiesse, Patrick Le Callet
MMSP3
2021 A machine-learning framework to predict TMO preference based on image and visual attention features
abstract
Tone-mapping operator (TMO) plays a crucial role in the task towards displaying high dynamic range (HDR) contents on standard displays. Similarly, visual attention (VA) is a vital feature of the human visual system (HVS), while it has not been sufficiently investigated in preference of tone-mapped images. The potential benefits of visual attention-based features for quality assessment of tone-mapped images are studied in this paper. A novel framework is proposed for tone-mapped image quality assessment to predict image preference. The framework is evaluated on two different datasets. Experimental results illustrate the importance of visual attention for improving the performance of objective metrics. The proposed framework outperforms the existing methods and presents as a competitive alternative for tone-mapped images evaluation.
Waqas Ellahi, Toinon Vigier, Patrick Le Callet
MMSP3
2021 Reliability of Crowdsourcing for Subjective Quality Evaluation of Tone Mapping Operators
abstract
Tone mapping operators (TMO) are functions which map high dynamic range (HDR) images to limited dynamic media while aiming to preserve the perceptual cues of the scene that govern its aesthetic quality. Evaluating aesthetic quality of TMOs is non-trivial due to the high subjectivity of preference involved. Traditionally, TMO aesthetic quality has been evaluated via subjective experiments in a controlled laboratory environment. However, the last decade has brought a surge in popularity of crowdsourcing as an alternative methodology to conduct subjective experiments. However, uncontrolled experiment conditions and unreliability of participant behaviour puts doubts on the trustworthiness of the collected data. In this study, we explore the possibility of using crowdsourcing platforms for subjective quality evaluation of TMOs. We have conducted three experiments with systematic changes to investigate the effect of experiment conditions and participant recruitment methods on the collected subjective data. Our results show that subjective evaluation of TMO aesthetic quality can be conducted on Prolific crowdsourcing platform with negligible differences in comparison to laboratory experiments. Furthermore, we provide objective conclusions about the effect of number of observers on the certainty of the pairwise comparison results.
Abhishek Goswami, Ali Ak, Wolf Hauser, Patrick Le Callet, Frédéric Dufaux
MMSP4
2021 Multi-Modal Aesthetic Assessment for Mobile Gaming Image
abstract
With the proliferation of various gaming technology, services, game styles, and platforms, multi-dimensional aesthetic assessment of the gaming contents is becoming more and more important for the gaming industry. Depending on the diverse needs of diversified game players, game designers, graphical developers, etc. in particular conditions, multi-modal aesthetic assessment is required to consider different aesthetic dimensions/perspectives. Since there are different underlying relationships between different aesthetic dimensions, e.g., between the ‘Colorfulness’ and ‘Color Harmony’, it could be advantageous to leverage effective information attached in multiple relevant dimensions. To this end, we solve this problem via multi-task learning. Our inclination is to seek and learn the correlations between different aesthetic relevant dimensions to further boost the generalization performance in predicting all the aesthetic dimensions. Therefore, the ‘bottleneck’ of obtaining good predictions with limited labeled data for one individual dimension could be unplugged by harnessing complementary sources of other dimensions, i.e., augment the training data indirectly by sharing training information across dimensions. According to experimental results, the proposed model outperforms state-of-the-art aesthetic metrics significantly in predicting four gaming aesthetic dimensions.
Yejing Xie, Suiyi Ling, Andreas Pastor, Junle Wang, Junyu Dong, Patrick Le Callet
MMSP7
2021 Exploring Crowdsourcing for Subjective Quality Assessment of 3D Graphics
abstract
Multimedia subjective quality assessment experiments are the most prominent and reliable way to evaluate the visual quality as perceived by human observers. Along with laboratory (lab) subjective experiments, crowdsourcing (CS) experiments have become very popular in recent years, e.g., during the COVID-19 pandemic these experiments provide an alternative to lab tests. However, conducting subjective quality assessment tests in CS raises many challenges: internet connection quality, lack of control on participants’ environment, participants’ consistency and reliability, etc. In this work, we evaluate the performance of CS studies for 3D graphics quality assessment To this end, we conducted a CS experiment based on the double stimulus impairment scale method and using a dataset of 80 meshes with diffuse color information corrupted by various distortions. We compared its results with those previously obtained in a lab study conducted on the same dataset and in a virtual reality environment. Results show that under controlled conditions and with appropriate participant screening strategies, a CS experiment can be as accurate as a lab experiment.
Yana Nehmé, Patrick Le Callet, Florent Dupont, Jean-Philippe Farrugia, Guillaume Lavoué
MMSP2
2021 The Effect of Temporal Sub-sampling on the Accuracy of Volumetric Video Quality Assessment
abstract
Volumetric video content has attracted increasing research interests over the last decade, as it facilitates the integration of dynamic real world content in virtual environments. Point cloud is one of the most common alternatives to represent volumetric video content. Yet, such representation requires an enormous data storage and pose significant greater pressures on compression algorithms compared to the standard 2D video. This challenge has unleashed a new wave in the development of novel point cloud compression technologies, which need to be evaluated in terms of production quality. Due to the high dimensionality of the data, evaluating the performances of relevant coding algorithms can be time consuming. This puts a barrier on optimizing coding algorithms with complex, but perceptually accurate, objective quality metrics. In this study, we thus explore the possibility of reducing temporal-dimension of the content under-evaluation, i.e., temporal sub-sampling, for objective quality evaluation without sacrificing from the correlation with the subjective opinion. In addition, we exploit different temporal pooling methods to further make the quality evaluation procedure more efficient. In total 30 different objective quality metrics were tested on the the V-SENSE volumetric video quality database. According to experimental results, there is no need to employ full frame-rate (30 fps) assessment to reach the meaningful correlation for the considered quality metrics. These observations could be referred to reduce the computation complexity regarding the evaluation and optimization of the relevant compression algorithms.
Ali Ak, Emin Zerman, Suiyi Ling, Patrick Le Callet, Aljoscha Smolic
PCS4
2021 Visualizing navigation difficulties in video game experiences
abstract
When developing video games, gameplay metrics allow to track and analyze the behaviors of users interacting with the game. Here, we propose to harness player's spatial trajectories to objectively quantity their gaming experience. Spatial trajectories are complex signals that are determined both by the players and by the topology of the virtual spaces they evolve in. In this paper, we propose a new methodology to measure and visualize how the entropy of trajectories is distributed in virtual spaces, and explain how it can inform game developers on the design of the game levels. We apply our method on the Sea Hero Quest dataset, consisting of the trajectories of over 4 millions players finding their way in water mazes.
Hippolyte Dubois, Patrick Le Callet, Antoine Coutrot
QoMEX2
2021 Ambiguity of objective image quality metrics: A new methodology for performance evaluation
Manri Cheon, Toinon Vigier, Lukas Krasula, Junghyuk Lee, Patrick Le Callet, Jong-Seok Lee
Signal Process. Image Commun.5
2021 Saliency4ASD: Challenge, dataset and tools for visual attention modeling for autism spectrum disorder
Jesús Gutiérrez 0001, Zhaohui Che, Guangtao Zhai, Patrick Le Callet
Signal Process. Image Commun.4
2021 Comparison of Subjective Methods for Quality Assessment of 3D Graphics in Virtual Reality
abstract
Numerous methodologies for subjective quality assessment exist in the field of image processing. In particular, the Absolute Category Rating with Hidden Reference (ACR-HR), the Double Stimulus Impairment Scale (DSIS), and the Subjective Assessment Methodology for Video Quality (SAMVIQ) are considered three of the most prominent methods for assessing the visual quality of 2D images and videos. Are these methods valid/accurate to evaluate the perceived quality of 3D graphics data? Is the presence of an explicit reference necessary, due to the lack of human prior knowledge on 3D graphics data compared to natural images/videos? To answer these questions, we compare these three subjective methods (ACR-HR, DSIS, and SAMVIQ) on a dataset of high-quality colored 3D models, impaired with various distortions. These subjective experiments were conducted in a virtual reality environment. Our results show differences in the performance of the methods depending on the 3D contents and the types of distortions. We show that DSIS and SAMVIQ outperform ACR-HR in terms of accuracy and point out a stable performance. In regard to the time-effort, DSIS achieves the highest accuracy in the shortest assessment time. Results also yield interesting conclusions on the importance of a reference for judging the quality of 3D graphics. We finally provide recommendations regarding the influence of the number of observers on the accuracy.
Yana Nehmé, Jean-Philippe Farrugia, Florent Dupont, Patrick Le Callet, Guillaume Lavoué
ACM Trans. Appl. Percept.4
2021 Perceptual Quality Assessment for Asymmetrically Distorted Stereoscopic Video by Temporal Binocular Rivalry
abstract
In this paper, we propose a two-stage weighting based perceptual quality assessment framework for asymmetrically distorted stereoscopic video (SV) sequences by temporal binocular rivalry. Firstly, a traditional 2D image quality assessment (IQA) method is employed to measure spatial distortion, and the temporal distortion is evaluated by the magnitude differences between motion vectors of distorted and reference video frames. Secondly, the structural strength (SS) computed by gradient map and the motion energy (ME) computed by frame difference map are used to estimate the intensity of visual stimulus in spatial and temporal domain respectively. Then, SS and ME are considered as the importance indexes to combine the quality scores of spatial and temporal distortion to estimate perceived distortion of single-view video sequences, which is denoted as the first-stage weighting. Finally, considering that the difference of intensity of visual stimulus between two eyes results in binocular rivalry, a novel temporal binocular rivalry inspired weighting method is designed to integrate the quality scores of left- and right-views for the final visual quality prediction of SV sequences, which is denoted as the second-stage weighting. Experimental results on Waterloo-IVC SV quality databases show that several specific examples of 2D-IQA methods within the proposed framework can obtain highly competitive performance over other existing ones.
Yuming Fang 0001, Xiangjie Sui, Jiheng Wang, Jiebin Yan, Jianjun Lei 0001, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.6
2021 Adversarial Attack Against Deep Saliency Models Powered by Non-Redundant Priors
abstract
Saliency detection is an effective front-end process to many security-related tasks, e.g. automatic drive and tracking. Adversarial attack serves as an efficient surrogate to evaluate the robustness of deep saliency models before they are deployed in real world. However, most of current adversarial attacks exploit the gradients spanning the entire image space to craft adversarial examples, ignoring the fact that natural images are high-dimensional and spatially over-redundant, thus causing expensive attack cost and poor perceptibility. To circumvent these issues, this paper builds an efficient bridge between the accessible partially-white-box source models and the unknown black-box target models. The proposed method includes two steps: 1) We design a new partially-white-box attack, which defines the cost function in the compact hidden space to punish a fraction of feature activations corresponding to the salient regions, instead of punishing every pixel spanning the entire dense output space. This partially-white-box attack reduces the redundancy of the adversarial perturbation. 2) We exploit the non-redundant perturbations from some source models as the prior cues, and use an iterative zeroth-order optimizer to compute the directional derivatives along the non-redundant prior directions, in order to estimate the actual gradient of the black-box target model. The non-redundant priors boost the update of some "critical" pixels locating at non-zero coordinates of the prior cues, while keeping other redundant pixels locating at the zero coordinates unaffected. Our method achieves the best tradeoff between attack ability and perturbation redundancy. Finally, we conduct a comprehensive experiment to test the robustness of 18 state-of-the-art deep saliency models against 16 malicious attacks, under both of white-box and black-box settings, which contributes a new robustness benchmark to the saliency community for the first time.
Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Yuan Tian 0017, Guodong Guo, Patrick Le Callet
IEEE Trans. Image Process.8
2021 Quality Assessment of Free-Viewpoint Videos by Quantifying the Elastic Changes of Multi-Scale Motion Trajectories
abstract
Virtual viewpoints synthesis is an essential process for many immersive applications including Free-viewpoint TV (FTV). A widely used technique for viewpoints synthesis is Depth-Image-Based-Rendering (DIBR) technique. However, such technique may introduce challenging non-uniform spatial-temporal structure-related distortions. Most of the existing state-of-the-art quality metrics fail to handle these distortions, especially the temporal structure inconsistencies observed during the switch of different viewpoints. To tackle this problem, an elastic metric and multi-scale trajectory based video quality metric (EM-VQM) is proposed in this paper. Dense motion trajectory is first used as a proxy for selecting temporal sensitive regions, where local geometric distortions might significantly diminish the perceived quality. Afterwards, the amount of temporal structure inconsistencies and unsmooth viewpoints transitions are quantified by calculating 1) the amount of motion trajectory deformations with elastic metric and, 2) the spatial-temporal structural dissimilarity. According to the comprehensive experimental results on two FTV video datasets, the proposed metric outperforms the state-of-the-art metrics designed for free-viewpoint videos significantly and achieves a gain of 12.86% and 16.75% in terms of median Pearson linear correlation coefficient values on the two datasets compared to the best one, respectively.
Suiyi Ling, Jing Li 0026, Zhaohui Che, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
IEEE Trans. Image Process.6
2021 Graph Learning Based Head Movement Prediction for Interactive 360 Video Streaming
abstract
Ultra-high definition (UHD) 360 videos encoded in fine quality are typically too large to stream in its entirety over bandwidth (BW)-constrained networks. One popular approach is to interactively extract and send a spatial sub-region corresponding to a viewer's current field-of-view (FoV) in a head-mounted display (HMD) for more BW-efficient streaming. Due to the non-negligible round-trip-time (RTT) delay between server and client, accurate head movement prediction foretelling a viewer's future FoVs is essential. In this paper, we cast the head movement prediction task as a sparse directed graph learning problem: three sources of relevant information-collected viewers' head movement traces, a 360 image saliency map, and a biological human head model-are distilled into a view transition Markov model. Specifically, we formulate a constrained maximum a posteriori (MAP) problem with likelihood and prior terms defined using the three information sources. We solve the MAP problem alternately using a hybrid iterative reweighted least square (IRLS) and Frank-Wolfe (FW) optimization strategy. In each FW iteration, a linear program (LP) is solved, whose runtime is reduced thanks to warm start initialization. Having estimated a Markov model from data, we employ it to optimize a tile-based 360 video streaming system. Extensive experiments show that our head movement prediction scheme noticeably outperformed existing proposals, and our optimized tile-based streaming scheme outperformed competitors in rate-distortion performance.
Xue Zhang 0008, Gene Cheung, Yao Zhao 0001, Patrick Le Callet, Chunyu Lin, Jack Z. G. Tan
IEEE Trans. Image Process.4
2021 Semi-Reference Sonar Image Quality Assessment Based on Task and Visual Perception
abstract
In submarine and underwater detection tasks, conventional optical imaging and analysis methods are not universally applicable due to the limited penetration depth of visible light. Instead, sonar imaging has become a preferred alternative. However, the capture and transmission conditions in complicated and dynamic underwater environments inevitably lead to visual quality degradation of sonar images, which might also impede further recognition, analysis and understanding. To measure this quality decrease and provide a solid quality indicator for sonar image enhancement, we propose a task- and perception-oriented sonar image quality assessment (TPSIQA) method, in which a semi-reference (SR) approach is applied to adapt to the limited bandwidth of underwater communication channels. In particular, we exploit reduced visual features that are critical for both human perception of and object recognition in sonar images. The final quality indicator is obtained through ensemble learning, which aggregates an optimal subset of multiple base learners to achieve both high accuracy and a high generalization ability. In this way, we are able to develop a compact but generalized quality metric using a small database of sonar images. Experimental results demonstrate competitive performance, high efficiency, and strong robustness of our method compared to the latest available image quality metrics.
Ke Gu 0001, Tiesong Zhao, Gangyi Jiang, Patrick Le Callet
IEEE Trans. Multim.5
2021 Wide Color Gamut Image Content Characterization: Method, Evaluation, and Applications
abstract
In this paper, we propose a novel framework to characterize a wide color gamut image content based on perceived quality due to the processes that change color gamut, and demonstrate two practical use cases where the framework can be applied. We first introduce the main framework and implementation details. Then, we provide analysis for understanding of existing wide color gamut datasets with quantitative characterization criteria on their characteristics, where four criteria, i.e., coverage, total coverage, uniformity, and total uniformity, are proposed. Finally, the framework is applied to content selection in a gamut mapping evaluation scenario in order to enhance reliability and robustness of the evaluation results. As a result, the framework fulfils content characterization for studies where quality of experience of wide color gamut stimuli is involved.
Junghyuk Lee, Toinon Vigier, Patrick Le Callet, Jong-Seok Lee
IEEE Trans. Multim.3
2021 Re-Visiting Discriminator for Blind Free-Viewpoint Image Quality Assessment
abstract
Accurate measurement of perceptual quality is important for various immersive multimedia, which demand real-time quality control or quality-based bench-marking for relevant algorithms. For instance, virtual views rendering in Free-Viewpoint (FV) navigation scenarios is a typical case that introduces challenging distortions, particularly the ones around dis-occluded regions. Existing quality metrics, most of which are targeting for impairments caused by compression or network condition, fail to quantify such non-uniform structure-related distortions. Moreover, the lack of quality databases for such distortions makes it even more challenging to develop robust quality metrics. In this work, a Generative Adversarial Networks based No-Reference (NR) quality Metric, namely GANs-NRM, is proposed. We first present an approach to create masks mimicking dis-occlusions/textureless regions, which is applicable on large-scale 2D image databases publicly available in the computer vision domain. Using these synthetic data, we then train a GANs-based context renderer with the capability of rendering those masked regions. Since the naturalness of the rendered dis-occluded regions strongly relates to the perceptual quality, we assume that the discriminator of the trained GANs has an intrinsic ability for quality assessment. We thus use the features extracted from the discriminator to learn a Bag-of-Distortion-Word (BDW) codebook. We show that a quality predictor can be then well trained using only a small amount of subjective quality data for the FV views rendering. Moreover, in the proposed framework, the discriminator is also adapted as a distortion-detector to locate possible distorted regions. According to the experimental results, the proposed model outperforms significantly the state-of-the-art quality metrics. The corresponding context renderer also shows appealing visualized results over other rendering algorithms.
Suiyi Ling, Jing Li 0026, Zhaohui Che, Wei Zhou 0021, Junle Wang, Patrick Le Callet
IEEE Trans. Multim.6
2021 Visual Quality of 3D Meshes With Diffuse Colors in Virtual Reality: Subjective and Objective Evaluation
abstract
Surface meshes associated with diffuse texture or color attributes are becoming popular multimedia contents. They provide a high degree of realism and allow six degrees of freedom (6DoF) interactions in immersive virtual reality environments. Just like other types of multimedia, 3D meshes are subject to a wide range of processing, e.g., simplification and compression, which result in a loss of quality of the final rendered scene. Thus, both subjective studies and objective metrics are needed to understand and predict this visual loss. In this work, we introduce a large dataset of 480 animated meshes with diffuse color information, and associated with perceived quality judgments. The stimuli were generated from 5 source models subjected to geometry and color distortions. Each stimulus was associated with 6 hypothetical rendering trajectories (HRTs): combinations of 3 viewpoints and 2 animations. A total of 11520 quality judgments (24 per stimulus) were acquired in a subjective experiment conducted in virtual reality. The results allowed us to explore the influence of source models, animations and viewpoints on both the quality scores and their confidence intervals. Based on these findings, we propose the first metric for quality assessment of 3D meshes with diffuse colors, which works entirely on the mesh domain. This metric incorporates perceptually-relevant curvature-based and color-based features. We evaluate its performance, as well as a number of Image Quality Metrics (IQMs), on two datasets: ours and a dataset of distorted textured meshes. Our metric demonstrates good results and a better stability than IQMs. Finally, we investigated how the knowledge of the viewpoint (i.e., the visible parts of the 3D model) may improve the results of objective metrics.
Yana Nehmé, Florent Dupont, Jean-Philippe Farrugia, Patrick Le Callet, Guillaume Lavoué
IEEE Trans. Vis. Comput. Graph.4
2020 A New Ensemble Adversarial Attack Powered by Long-Term Gradient Memories
abstract
Deep neural networks are vulnerable to adversarial attacks. More importantly, some adversarial examples crafted against an ensemble of pre-trained source models can transfer to other new target models, thus pose a security threat to black-box applications (when the attackers have no access to the target models). Despite adopting diverse architectures and parameters, source and target models often share similar decision boundaries. Therefore, if an adversary is capable of fooling several source models concurrently, it can potentially capture intrinsic transferable adversarial information that may allow it to fool a broad class of other black-box target models. Current ensemble attacks, however, only consider a limited number of source models to craft an adversary, and obtain poor transferability. In this paper, we propose a novel black-box attack, dubbed Serial-Mini-Batch-Ensemble-Attack (SMBEA). SMBEA divides a large number of pre-trained source models into several mini-batches. For each single batch, we design 3 new ensemble strategies to improve the intra-batch transferability. Besides, we propose a new algorithm that recursively accumulates the “long-term” gradient memories of the previous batch to the following batch. This way, the learned adversarial information can be preserved and the inter-batch transferability can be improved. Experiments indicate that our method outperforms state-of-the-art ensemble attacks over multiple pixel-to-pixel vision tasks including image translation and salient region prediction. Our method successfully fools two online black-box saliency prediction systems including DeepGaze-II (Kummerer 2017) and SALICON (Huang et al. 2017). Finally, we also contribute a new repository to promote the research on adversarial attack and defense over pixel-to-pixel tasks: https://github.com/CZHQuality/AAA-Pix2pix.
Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Patrick Le Callet
AAAI6
2020 Few-Shot Pill Recognition
abstract
Pill image recognition is vital for many personal/public health-care applications and should be robust to diverse unconstrained real-world conditions. Most existing pill recognition models are limited in tackling this challenging few-shot learning problem due to the insufficient instances per category. With limited training data, neural network-based models have limitations in discovering most discriminating features, or going deeper. Especially, existing models fail to handle the hard samples taken under less controlled imaging conditions. In this study, a new pill image database, namely CURE, is first developed with more varied imaging conditions and instances for each pill category. Secondly, a W2-net is proposed for better pill segmentation. Thirdly, a Multi-Stream (MS) deep network that captures task-related features along with a novel two-stage training methodology are proposed. Within the proposed framework, a Batch All strategy that considers all the samples is first employed for the sub-streams, and then a Batch Hard strategy that considers only the hard samples mined in the first stage is utilized for the fusion network. By doing so, complex samples that could not be represented by one type of feature could be focused and the model could be forced to exploit other domain-related information more effectively. Experiment results show that the proposed model outperforms state-of-the-art models on both the National Institute of Health (NIH) and our CURE database.
Suiyi Ling, Andreas Pastor, Jing Li 0026, Zhaohui Che, Junle Wang, Patrick Le Callet
CVPR7
2020 Sparse Directed Graph Learning for Head Movement Prediction in 360 Video Streaming
abstract
High-definition 360 videos encoded in fine quality are typically too large in size to stream in its entirety over bandwidth (BW)-constrained networks. One popular remedy is to interactively extract and send a spatial sub-region corresponding to a viewer's current field-of-view (FoV) in a head-mounted display (HMD) for more BW-efficient streaming. Due to the non-negligible round-trip-time (RTT) delay between server and client, accurate head movement prediction that foretells a viewer's future FoVs is essential. Existing approaches are either overly simplistic in modelling and predict poorly when RTT is large, or are over-reliant on data-driven learning, resulting in inflexible models that are not robust to RTT heterogeneity. In this paper, we cast the head movement prediction task as a sparse directed graph learning problem, where three sources of relevant information-a 360 image saliency map, collected viewers' head movement traces, and a biological head rotation model-are aggregated into a unified Markov model. Specifically, we formulate a constrained optimization problem to minimize an l2-norm fidelity term and a sparsity term, corresponding to trace data / saliency consistency and a sparse graph model prior respectively. We solve the problem alternately using a hybrid iterative reweighted least square (IRLS) and Frank-Wolfe optimization strategy. Extensive experiments show that our head movement prediction scheme noticeably outperforms existing proposals across a wide range of RTTs.
Xue Zhang 0008, Gene Cheung, Patrick Le Callet, Jack Z. G. Tan
ICASSP3
2020 Towards Visual Saliency Computation on 3D Graphical Contents for Interactive Visualization
abstract
Understanding human visual attention mechanisms and interaction in immersive scenes are of great importance in perception. In immersive context, users are able to interact with increasingly rich/ complex 3D contents during rendering. Therefore, to avoid latency or rendering issues, there is a critical need for simplifying and filtering the primitives and levels of detail of these high-quality 3D graphics (according to viewing conditions). In order to ensure a high user's quality of experience (QoE) during interactive visualization, these processing operations should take into account perceptual information. To do so, we suggest an approach that uses visual saliency information of the 3D scene to guide simplification and level of details selection. In this paper, we question the efficiency of our novel approach to compute visual saliency on 3D graphics. This approach takes into consideration the viewpoint from which the 3D content was seen/rendered when computing saliency (whichever the considered viewpoint is), by using saliency maps of view-based method computed offline. Such technique could help alleviate rendering constraints during interactive visualization.
Mona Abid, Matthieu Perreira Da Silva, Patrick Le Callet
ICIP3
2020 A Case Study of Machine Learning Classifiers for Real-Time Adaptive Resolution Prediction in Video Coding
abstract
At lower bit-rate encoding video in real-time with a reasonable viewing quality is challenging. Content adaptive per-title encoding is usually leveraged for OTT/VOD delivery by selecting the optimal resolutions and qualities of a given video using multiple encodings. Built on such powerful resolution selection principles, this paper introduces an on the fly resolution prediction without requiring multiple encoding with the help of machine learning which is suitable for real-time video delivery. Two machine learning networks are defined based on the resolution of the previous decision period. Three types of machine learning classifiers: weighted SVM, Random Forests (RF), and custom-designed Multi-Layer Perceptron (MLP) are tested. Suitability of classifiers for real-time resolution prediction is discussed based on the accuracy, BDrate performances, and impact of misclassification on encoding performance and hardware implementability. The proposed solution offers a promising average bit-rate savings upto 12.6%.
Madhukar Bhat, Jean-Marc Thiesse, Patrick Le Callet
ICME3
2020 Towards Perceptually-Optimized Compression Of User Generated Content (UGC): Prediction Of UGC Rate-Distortion Category
abstract
How to best evaluate the perceptual quality, and efficiently optimize the compression of User Generated Content (UGC) within an adaptive streaming system is becoming one of the most intractable challenges in the community. Rate-Distortion (R-D) characteristic based content analyses, which could be applied on the non-pristine originals, is inevitable to provide guidance in developing quality metrics and efficient compression system. To this end, we present a novel complete R-D category prediction system through the identification of discriminate features. To better understand the Rate-Distortion (R-D) behaviors of UGC, we first propose a Bjontegaard Delta (BD)-Rate, BD-Quality-based algorithm to categorize UGC. By using the predicted R-D related categories as ground-truth labels, we further identify features that characterize the R-D behaviors of UGC via a hierarchical feature selection framework. Finally, selected features are employed to predict the R-D category of under-test UGC. Comprehensive observations and results are summarized through extensive experiments.
Suiyi Ling, Yoann Baveye, Patrick Le Callet, Jim Skinner, Ioannis Katsavounidis
ICME3
2020 A Probabilistic Graphical Model for Analyzing the Subjective Visual Quality Assessment Data from Crowdsourcing
abstract
The swift development of the multimedia technology has raised dramatically the users' expectation on the quality of experience. To obtain the ground-truth perceptual quality for model training, subjective assessment is necessary. Crowdsourcing platform provides us a convenient and feasible way to run large-scale experiments. However, the obtained perceptual quality labels are generally noisy. In this paper, we propose a probabilistic graphical annotation model to infer the underlying ground truth and discovering the annotator's behavior. In the proposed model, the ground truth quality label is considered following a categorical distribution rather than a unique number, i.e., different reliable opinions on the perceptual quality are allowed. In addition, different annotator's behaviors in crowdsourcing are modeled, which allows us to identify the possibility that the annotator makes noisy labels during the test. The proposed model has been tested on both simulated data and real-world data, where it always shows superior performance than the other state-of-the-art models in terms of accuracy and robustness.
Jing Li 0026, Suiyi Ling, Junle Wang, Patrick Le Callet
ACM Multimedia4
2020 QoEVMA'20: 1st Workshop on Quality of Experience (QoE) in Visual Multimedia Applications
abstract
Nowadays, people spend dramatically more time on watching videos through different devices. The advanced hardware technology and network allow for the increasing demands of users viewing experience. Thus, enhancing the Quality of Experience of end-users in advanced multimedia is the ultimate goal of service providers, as good services would attract more consumers. Quality assessment is thus important. The first workshop on "Quality of Experience (QoE) in visual multimedia applications" (QoEVMA'20) focuses on the QoE assessment of any visual multimedia applications both subjectively and objectively. The topics include 1)QoE assessment on different visual multimedia applications, including VoD for movies, dramas, variety shows, UGC on social networks, live streaming videos for gaming/shopping/social, etc. 2)QoE assessment for different video formats in multimedia services, including 2D, stereoscopic 3D, High Dynamic Range (HDR), Augmented Reality (AR), Virtual Reality (VR), 360, Free-Viewpoint Video(FVV), etc. 3)Key performance indicators (KPI) analysis for QoE. This summary gives a brief overview of the workshop, which took place at October 16, 2020 in Seattle (U.S.), as a half-day workshop.
Xinbo Gao 0001, Patrick Le Callet, Jing Li 0026, Zhi Li 0001, Wen Lu 0004
ACM Multimedia2
2020 Few-Shot Object Detection in Real Life: Case Study on Auto-Harvest
abstract
Confinement during COVID-19 has caused serious effects on agriculture all over the world. As one of the efficient solutions, mechanical harvest/auto-harvest that is based on object detection and robotic harvester becomes an urgent need. Within the auto-harvest system, robust few-shot object detection model is one of the bottlenecks, since the system is required to deal with new vegetable/fruit categories and the collection of large-scale annotated datasets for all the novel categories is expensive. There are many few-shot object detection models that were developed by the community. Yet whether they could be employed directly for real life agricultural applications is still questionable, as there is a context-gap between the commonly used training datasets and the images collected in real life agricultural scenarioas. To this end, in this study, we present a novel cucumber dataset and propose two data augmentation strategies that help to bridge the context-gap. Experimental results show that 1) the state-of-the-art few-shot object detection model performs poorly on the novel `cucumber' category; and 2) the proposed augmentation strategies outperform the commonly used ones.
Kévin Riou, Suiyi Ling, Mathis Piquet, Vincent Truffault, Patrick Le Callet
MMSP6
2020 Can Visual Scanpath Reveal Personal Image Memorability? Investigation of HMM Tools for Gaze Patterns Analysis
abstract
Visual attention has been shown as a good proxy for QoE, revealing specific visual patterns considering content, system and contextual aspects of a multimedia applications. In this paper, we propose a novel approach based on hidden markov models to analyze visual scanpaths in an image memorability task. This new method ensures the consideration of both temporal and idiosyncrasic aspects of visual behavior. The study shows promising results for the use of indirect measures for the personalization of QoE assessment and prediction.
Waqas Ellahi, Toinon Vigier, Patrick Le Callet
QoMEX3
2020 Modelling effects of S3D visual discomfort in human emotional state using data mining techniques
Dragana Dordevic Cegar, Miguel Barreda-Ángeles, Dragan Kukolj, Patrick Le Callet
Multim. Tools Appl.4
2020 How is Gaze Influenced by Image Transformations? Dataset and Model
abstract
Data size is the bottleneck for developing deep saliency models, because collecting eye-movement data is very time-consuming and expensive. Most of current studies on human attention and saliency modeling have used high-quality stereotype stimuli. In real world, however, captured images undergo various types of transformations. Can we use these transformations to augment existing saliency datasets? Here, we first create a novel saliency dataset including fixations of 10 observers over 1900 images degraded by 19 types of transformations. Second, by analyzing eye movements, we find that observers look at different locations over transformed versus original images. Third, we utilize the new data over transformed images, called data augmentation transformation (DAT), to train deep saliency models. We find that label-preserving DATs with negligible impact on human gaze boost saliency prediction, whereas some other DATs that severely impact human gaze degrade the performance. These label-preserving valid augmentation transformations provide a solution to enlarge existing saliency datasets. Finally, we introduce a novel saliency model based on generative adversarial networks (dubbed GazeGAN). A modified U-Net is utilized as the generator of the GazeGAN, which combines classic "skip connection" with a novel "center-surround connection" (CSC) module. Our proposed CSC module mitigates trivial artifacts while emphasizing semantic salient regions, and increases model nonlinearity, thus demonstrating better robustness against transformations. Extensive experiments and comparisons indicate that GazeGAN achieves state-of-the-art performance over multiple datasets. We also provide a comprehensive comparison of 22 saliency models on various transformed scenes, which contributes a new robustness benchmark to saliency community. Our code and dataset are available at.
Zhaohui Che, Ali Borji, Guangtao Zhai, Xiongkuo Min, Guodong Guo, Patrick Le Callet
IEEE Trans. Image Process.6
2020 A Metric for Light Field Reconstruction, Compression, and Display Quality Evaluation
abstract
Owning to the recorded light ray distributions, light field contains much richer information and provides possibilities of some enlightening applications, and it has becoming more and more popular. To facilitate the relevant applications, many light field processing techniques have been proposed recently. These operations also bring the loss of visual quality, and thus there is need of a light field quality metric to quantify the visual quality loss. To reduce the processing complexity and resource consumption, light fields are generally sparsely sampled, compressed, and finally reconstructed and displayed to the users. We consider the distortions introduced in this typical light field processing chain, and propose a full-reference light field quality metric. Specifically, we measure the light field quality from three aspects: global spatial quality based on view structure matching, local spatial quality based on near-edge mean square error, and angular quality based on multi-view quality analysis. These three aspects have captured the most common distortions introduced in light field processing, including global distortions like blur and blocking, local geometric distortions like ghosting and stretching, and angular distortions like flickering and sampling. Experimental results show that the proposed method can estimate light field quality accurately, and it outperforms the state-of-the-art quality metrics which may be effective for light field.
Xiongkuo Min, Jiantao Zhou 0001, Guangtao Zhai, Patrick Le Callet, Xiaokang Yang 0001, Xin-Ping Guan
IEEE Trans. Image Process.4
2020 Training Objective Image and Video Quality Estimators Using Multiple Databases
abstract
Machine learning (ML) is an essential part of recent advances in computer science. To fully exploit its potential, ML-based algorithms require a considerable amount of annotated data to be used for training. This represents a severe limitation in the field of image and video quality assessment since obtaining large-scale annotated databases is time-consuming and expensive. Moreover, the resulting quality estimators are mainly restricted only to the usecases included in the dataset used for their training. This paper proposes a strategy allowing for combination of multiple databases for training of objective image and video quality assessment algorithms. Using this strategy, the algorithms can be trained using all of the existing relevant databases together which allows to increase the amount of data-points and usecases in orders of magnitude. The potential of the proposed method is demonstrated by re-training the combination of features from Video Multimethod Assessment Fusion (VMAF) algorithm resulting in the significant improvement of its performance with respect to 20 video databases.
Lukas Krasula, Yoann Baveye, Patrick Le Callet
IEEE Trans. Multim.3
2020 FFTMI: Features Fusion for Natural Tone-Mapped Images Quality Evaluation
abstract
Tone-mapping is a crucial step in the task towards displaying high dynamic range (HDR) images on standard displays. Given the number of possible ways to tone-map such images, development of an objective quality criterion, enabling selection of the most suitable tone-mapping operator (TMO) and setting its parameters in order to maximize the quality of the reproduction, is of high interest. In this paper, a new objective metric for natural tone-mapped images is proposed. It is based on a fusion of several perceptually relevant features that have been carefully selected using an appropriate feature selection procedure. The outcome of the selection also provides a valuable insight into the importance of particular perceptual aspects when judging the quality of tone-mapped HDR content. The performance of the resulting combination of features is thoroughly evaluated with respect to three publicly available databases and compared to several relevant state-of-the-art criteria. The proposed approach is shown to significantly outperform the tested metrics and can, therefore, be considered a competitive alternative for tone-mapped images evaluation.
Lukas Krasula, Karel Fliegel, Patrick Le Callet
IEEE Trans. Multim.3
2019 Comparison of subjective methods, with and without explicit reference, for quality assessment of 3D graphics
abstract
Numerous methodologies for subjective quality assessment exist in the field of image processing. In particular, the Absolute Category Rating with Hidden Reference (ACR-HR) and the Double Stimulus Impairment Scale (DSIS) are considered two of the most prominent methods for assessing the visual quality of 2D images and videos. Are these methods valid/accurate to evaluate the perceived quality of 3D graphics data? Is the presence of an explicit reference necessary, due to the lack of human prior knowledge on 3D graphics data compared to natural images/videos? To answer these questions, we compare these two subjective methods (ACR-HR and DSIS) on a dataset of high-quality colored 3D models, impaired with various distortions. These subjective experiments were conducted in a virtual reality (VR) environment. Our results show differences in the performance of the methods depending on the 3D contents and the types of distortions. We show that DSIS outperforms ACR-HR in term of accuracy and points out a stable performance. Results also yield interesting conclusions on the importance of a reference for judging the quality of 3D graphics. We finally provide recommendations regarding the influence of the number of observers on the accuracy.
Yana Nehmé, Jean-Philippe Farrugia, Florent Dupont, Patrick Le Callet, Guillaume Lavoué
SAP4
2019 Influence of Viewpoint on Visual Saliency Models for Volumetric Content
abstract
In order to predict where humans look in a 3D immersive environment, saliency can be computed using either 3D saliency models or view-based approaches (2D projection). In fact, building a 3D complete model is still a challenging task that is not investigated enough in the research field while 2D imaging approaches have been extensively studied and have shown solid performances.As 6 degrees of freedom are allowed in volumetric videos, users are able to navigate through the content in different manners. In this case, 2D saliency models might be less robust if applied naïvely, since advanced parameters such as viewing distance are not considered in such models.The aim of this paper is to investigate the influence of viewpoint on 2D saliency models when applied on volumetric data and this to get a better understanding of how viewpoint information could be integrated into view-based approaches.To do so, a subjective psycho-visual experiment was conducted and a fine analysis was led using the variance analysis statistical method.
Mona Abid, Matthieu Perreira Da Silva, Patrick Le Callet
ICIP3
2019 On Accuracy of Objective Metrics for Assessment of Perceptual Pre-Processing for Video Coding
abstract
The objective assessment of encoding performance is a key aspect of video delivery optimization. Objective metrics typically do not address fully the different viewing distances and behavior of compression artifacts being subjected to perceptual changes in video. This poses a daunting task of opti-mizing compression for video delivery systems for specific viewing conditions and perceptual optimization. This paper introduces a perceptual pre-processing and then discusses accuracy of typically used objective metrics for judging performance of pre-processing for observers at different viewing distances. To this end, this paper reports an in-depth analysis to check if objective metrics can successfully match subjective results at critical pairs i.e. pre-processed and original video at same QP.
Madhukar Bhat, Jean-Marc Thiesse, Patrick Le Callet
ICIP3
2019 Perceptual Representations of Structural Information in Images: Application to Quality Assessment of Synthesized View in FTV Scenario
abstract
As the immersive multimedia techniques like Free-viewpoint TV (FTV) develop at an astonishing rate, user's demand for high-quality immersive contents increases dramatically. Unlike traditional uniform artifacts, the distortions within immersive contents could be non-uniform structure-related and thus are challenging for commonly used quality metrics. Recent studies have demonstrated that the representation of visual features can be extracted from multiple levels of the hierarchy. Inspired by the hierarchical representation mechanism in the human visual system (HVS), in this paper, we explore to adopt structural representations to quantitatively measure the impact of such structure-related distortion on perceived quality in FTV scenario. More specifically, a bio-inspired full reference image quality metric is proposed based on 1) low-level contour descriptor; 2) mid-level contour category descriptor; and 3) task-oriented non-natural structure descriptor. The experimental results show that the proposed model outperforms significantly the state-of-the-art metrics.
Suiyi Ling, Jing Li 0026, Patrick Le Callet, Junle Wang
ICIP3
2019 On the usage of visual saliency models for computer generated objects
abstract
Visual attention is a key feature to optimize visual experience of many multimedia applications. 2D visual attention computational modeling is an active research area considering the visualization of natural images on a conventional display. In this paper, we question the ability of such models to be applicable to single computer-generated objects rendered at different sizes (on a conventional display). We benchmark state of art visual attention models and investigate the influence of the viewpoint on those computational models applied on volumetric data and this to get a better understanding of how viewpoint information could be integrated into view-based approaches. To do so, a subjective experiment was conducted and a fine analysis was led using the variance analysis statistical method.
Mona Abid, Matthieu Perreira Da Silva, Patrick Le Callet
MMSP3
2019 A dataset of eye movements for the children with autism spectrum disorder
abstract
Social difficulties are the hallmark features of Autism Spectrum Disorder (ASD) and can lead to atypical visual attention towards stimuli. Eye movements encode rich information about attention and psychological factors of an individual, which could help to characterize the traits of ASD. Learning atypical eye movements of the individuals with ASD towards various stimuli is important and has many application scenarios. However, due to the lack of open datasets, research in this sense is still limited. In this work, we present an open dataset of eye movements of children with Autism Spectrum Disorder. It consists of 300 natural scene images and the corresponding eye movement data collected from 14 children with ASD and 14 healthy controls. In particular, fixation maps and scanpaths are available in the dataset. Based on this dataset, researchers could analyze the visual traits of children with ASD and design specialized visual attention models to promote research in related fields, as well as design specialized models to identify the individuals with ASD. The dataset can be accessed in http://doi.org/10.5281/zenodo.2647418
Huiyu Duan, Guangtao Zhai, Xiongkuo Min, Zhaohui Che, Yi Fang 0009, Xiaokang Yang 0001, Jesús Gutiérrez 0001, Patrick Le Callet
MMSys8
2019 Contactless approach for heart rate estimation for QoE assessment
Mattia Bonomi, Federica Battisti, Giulia Boato, Miguel Barreda-Ángeles, Marco Carli, Patrick Le Callet
Signal Process. Image Commun.6
2019 Quality assessment for view synthesis using low-level and mid-level structural representation
Yu Zhou 0009, Leida Li, Suiyi Ling, Patrick Le Callet
Signal Process. Image Commun.4
2019 Reference-Free Quality Assessment of Sonar Images via Contour Degradation Measurement
abstract
Sonar imagery plays a significant role in oceanic applications since there is little natural light underwater, and light is irrelevant to sonar imaging. Sonar images are very likely to be affected by various distortions during the process of transmission via the underwater acoustic channel for further analysis. At the receiving end, the reference image is unavailable due to the complex and changing underwater environment and our unfamiliarity with it. To the best of our knowledge, one of the important usages of sonar images is target recognition on the basis of contour information. The contour degradation degree for a sonar image is relevant to the distortions contained in it. To this end, we developed a new no-reference contour degradation measurement for perceiving the quality of sonar images. The sparsities of a series of transform coefficient matrices, which are descriptive of contour information, are first extracted as features from the frequency and spatial domains. The contour degradation degree for a sonar image is then measured by calculating the ratios of extracted features before and after filtering this sonar image. Finally, a bootstrap aggregating (bagging)-based support vector regression module is learned to capture the relationship between the contour degradation degree and the sonar image quality. The results of experiments validate that the proposed metric is competitive with the state-of-the-art reference-based quality metrics and outperforms the latest reference-free competitors.
Ke Gu 0001, Weisi Lin, Zhifang Xia, Patrick Le Callet, En Cheng
IEEE Trans. Image Process.5
2019 Fast Blind Quality Assessment of DIBR-Synthesized Video Based on High-High Wavelet Subband
abstract
Free-viewpoint video, as the development direction of the next-generation video technologies, uses the depth-image-based rendering (DIBR) technique for the synthesis of video sequences at viewpoints, where real captured videos are missing. As reference videos at multiple viewpoints are not available, a blind reliable real-time quality metric of the synthesized video is needed. Although no-reference quality metrics dedicated to synthesized views successfully evaluate synthesized images, they are not that effective when evaluating synthesized video due to additional temporal flicker distortion typical only for video. In this paper, a new fast no-reference quality metric of synthesized video with synthesis distortions is proposed. It is guided by the fact that the DIBR-synthesized images are characterized by increased high frequency content. The metric is designed under the assumption that the perceived quality of DIBR-synthesized video can be estimated by quantifying the selected areas in the high-high wavelet subband. The threshold is used to select the most important distortion sensitive regions. The proposed No-Reference Morphological Wavelet with Threshold (NR_MWT) metric is computationally extremely efficient, comparable to PSNR, as the morphological wavelet transformation uses very short filters and only integer arithmetic. It is completely blind, without using machine learning techniques. Tested on the publicly available dataset of synthesized video sequences characterized by synthesis distortions, the metric achieves better performances and higher computational efficiency than the state-of-the-art metrics dedicated to DIBR-synthesized images and videos.
Dragana Sandic-Stankovic, Dragan Kukolj, Patrick Le Callet
IEEE Trans. Image Process.3
2019 Improved Performance Measures for Video Quality Assessment Algorithms Using Training and Validation Sets
abstract
The training and performance analysis of objective video quality assessment algorithms is complex due to the huge variety of possible content classes and transmission distortions. Several secondary issues such as free parameters in machine learning algorithms and alignment of subjective datasets put an additional burden on the developer. In this paper, three subsequent steps are presented to address such issues. First, the content and coding parameter space of a large-scale database is used to select dedicated subsets for training objective algorithms. This aims at providing a method for selecting the most significant contents and coding parameters from all imaginable combinations. In the practical case where only a limited set is available, it also helps us to avoid redundancy in the training subset selection. The second step is a discussion on performance measures for algorithms that employ machine-learning methods. The particularity of the performance measures is that the quality of the training and verification datasets is taken into consideration. Common issues that often use existing measures are presented, and improved or complementary methods are proposed. The measures are applied to two examples of no-reference objective assessment algorithms using the aforementioned subsets of the large-scale database. While limited in terms of practical applications, this sandbox approach of objectively predicting an objectively evaluated video sequences allows for eliminating additional influence factors from subjective studies. In the third step, the proposed performance measures are applied to the practical case of training and analyzing assessment algorithms on readily available subjectively annotated image datasets. The presentation method in this part of the paper can also be used as an exemplified recommendation for reporting in-depth information on the performance. Using this presentation method, future publications presenting newly developed quality assessment algorithms may be significantly improved.
Ahmed Aldahdooh, Enrico Masala, Olivier Janssens, Glenn Van Wallendael, Marcus Barkowsky, Patrick Le Callet
IEEE Trans. Multim.6
2019 AccAnn: A New Subjective Assessment Methodology for Measuring Acceptability and Annoyance of Quality of Experience
abstract
User expectations have a crucial impact on the levels of quality of experience (QoE) that they consider acceptable or satisfying. Measuring acceptability and annoyance has mainly been performed in separate or multi-step experiments without any control over participants' expectations. This paper introduces a simple methodology to obtain the information about both of the entities in a single step and compares several data processing strategies useful for results interpretation. A specifically designed subjective experiment, conducted on compressed videos, has shown that the multi-step procedures could be replaced by our proposed single-step approach, regardless of the viewing conditions, while the novel approach is significantly preferred by observers for its low time requirements and higher intuitiveness. The test has simultaneously proven that user expectations can be altered by the instructions and it is, therefore, possible to simulate different user profiles regardless of the participants' real habits. The acceptability/annoyance experimental results are also used to benchmark the state-of-the-art objective video quality metrics in predicting acceptability/annoyance of QoE. A case study on the determination of the threshold of acceptability/annoyance for objective quality metrics is conducted, which can be served as a guideline for video streaming service providers.
Jing Li 0026, Lukas Krasula, Yoann Baveye, Zhi Li 0001, Patrick Le Callet
IEEE Trans. Multim.5
2018 A Blind Quality Measure for Industrial 2D Matrix Symbols Using Shallow Convolutional Neural Network
abstract
Industrial two-dimensional (2D) matrix symbols are ubiquitous throughout the automatic assembly lines. Most industrial 2D symbols are corrupted by various inevitable artifacts. State-of-the-art decoding algorithms are not able to directly handle low-quality symbols irrespective of problematic artifacts. Degraded symbols require appropriate preprocessing methods, such as morphology filtering, median filtering, or sharpening filtering, according to specific distortion type. In this paper, we first establish a database including 3000 industrial 2D symbols which are degraded by 6 types of distortions. Second, we utilize a shallow convolutional neural network (CNN) to identify the distortion type and estimate the quality grade for 2D symbols. Finally, we recommend an appropriate preprocessing method for low-quality symbol according to its distortion type and quality grade. Experimental results indicate that the proposed method outperforms state-of-the-art methods in terms of PLCC, SRCC and RMSE. It also promotes decoding efficiency at the cost of low extra time spent.
Zhaohui Che, Guangtao Zhai, Jing Liu 0002, Ke Gu 0001, Patrick Le Callet, Jiantao Zhou 0001, Xianming Liu 0005
ICIP5
2018 A Perception-Based Framework for Wide Color Gamut Content Selection
abstract
Considering the content dependence of the perceived quality, selection of source content can significantly influence the results of studies related to Quality of Experience. In this paper, we propose an automated content selection method towards wide color gamut stimuli. The framework enables to objectively characterize the content according to its perceptual properties and thus allows to select a representative, diverse, and challenging subsets for various studies. Experimental results validate the reliability and robustness of the proposed framework.
Junghyuk Lee, Toinon Vigier, Patrick Le Callet, Jong-Seok Lee
ICIP3
2018 How to Learn the Effect of Non-Uniform Distortion on Perceived Visual Quality? Case Study Using Convolutional Sparse Coding for Quality Assessment of Synthesized Views
abstract
Machine learning has been attached greater attention in the field of quality assessment and related models have been developed by assuming that any region within images share the same quality score. However, this assumption may no longer stand when non-uniform distortions exist, e.g. geometric distortion within synthesized views in the case of free- view point video TV (FTV). In this paper, we explore to use Convolutional Sparse Coding (CSC), which computes a sparse representation for an entire image with the sum of a set of convolutions with dictionary filters instead of a linear combination of a set of dictionary atoms, to learn a visible codebook and propose a methodology to quantify the geometric distortion without using the reference. Experimental results show that the proposed no reference metric is the most visual friendly and reliable metric among the compared blind image metrics designed for synthesized images.
Suiyi Ling, Patrick Le Callet
ICIP2
2018 No-Reference Quality Assessment for Stitched Panoramic Images Using Convolutional Sparse Coding and Compound Feature Selection
abstract
Image stitching-composition of different viewpoint images to form a 360-degree panoramic image-is an essential component towards immersive applications like VR and AR. A no-reference (NR) quality metric specifically designed to evaluate stitched panoramic images is highly desirable when ground-truth reference images are not available. In this paper, we use Convolutional Sparse Coding (CSC) with a set of convolutional filters to locate stitching-specific distortions in a target image, and design trained kernels to quantify the compound effects of multiple distortion types in a local region. Specifically, our contributions are: i) a training database labeled with location information of the distortion regions is released; ii) a NR metric is proposed to accurately assess stitching-specific artifacts like ghosting using convolutional sparse coding; and iii) a novel sequential feature selection algorithm is proposed to quantify the aforementioned compound distortion effects. In extensive experiments, we show that the performance of our proposed NR metric is comparable to the state-of-the-art full-reference metrics designed for stitched images.
Suiyi Ling, Gene Cheung, Patrick Le Callet
ICME3
2018 Adaptive Screen Content Image Enhancement Strategy using Layer-based Segmentation
abstract
The ubiquitous screen content images (SCIs) play a significant role in various scenarios currently. However, most SCIs captured by consumer devices are frequently corrupted with distortions, especially contrast distortion. Unlike the natural images, SCIs are composed of text, graphics and natural scene pictures so that traditional image enhancement methods are not suitable for these compound images. Therefore, we innovatively proposed an adaptive strategy for enhancing SCIs in this paper. Firstly, we devised a segmentation method to divide SCI into text and pictorial regions. Next, the famous guided image filter (GIF) with big and small kernel sizes served as unsharpness masking for processing different regions adaptively. For verifying performance, the proposed method was tested on recently prevalent SCI datasets including SIQAD, and Webpage Dataset. Experimental results indicate that the proposed approach outperforms state-of-the-art methods in most SCIs with flat background.
Zhaohui Che, Guangtao Zhai, Ke Gu 0001, Patrick Le Callet, Xianming Liu 0005, Deming Zhai, Xiao Gu 0001
ISCAS4
2018 A dataset of head and eye movements for 360° videos
abstract
Research on visual attention in 360° content is crucial to understand how people perceive and interact with this immersive type of content and to develop efficient techniques for processing, encoding, delivering and rendering. And also to offer a high quality of experience to end users. The availability of public datasets is essential to support and facilitate research activities of the community. Recently, some studies have been presented analyzing exploration behaviors of people watching 360° videos, and a few datasets have been published. However, the majority of these works only consider head movements as proxy for gaze data, despite the importance of eye movements in the exploration of omnidirectional content. Thus, this paper presents a novel dataset of 360° videos with associated eye and head movement data, which is a follow-up to our previous dataset for still images [14]. Head and eye tracking data was obtained from 57 participants during a free-viewing experiment with 19 videos. In addition, guidelines on how to obtain saliency maps and scanpaths from raw data are provided. Also, some statistics related to exploration behaviors are presented, such as the impact of the longitudinal starting position when watching omnidirectional videos was investigated in this test. This dataset and its associated code are made publicly available to support research on visual attention for 360° content.
Erwan J. David, Jesús Gutiérrez 0001, Antoine Coutrot, Matthieu Perreira Da Silva, Patrick Le Callet
MMSys5
2018 Hybrid-MST: A Hybrid Active Sampling Strategy for Pairwise Preference Aggregation
abstract
In this paper we present a hybrid active sampling strategy for pairwise preference aggregation, which aims at recovering the underlying rating of the test candidates from sparse and noisy pairwise labeling. Our method employs Bayesian optimization framework and Bradley-Terry model to construct the utility function, then to obtain the Expected Information Gain (EIG) of each pair. For computational efficiency, Gaussian-Hermite quadrature is used for estimation of EIG. In this work, a hybrid active sampling strategy is proposed, either using Global Maximum (GM) EIG sampling or Minimum Spanning Tree (MST) sampling in each trial, which is determined by the test budget. The proposed method has been validated on both simulated and real-world datasets, where it shows higher preference aggregation ability than the state-of-the-art methods.
Jing Li 0026, Rafal Mantiuk, Junle Wang, Suiyi Ling, Patrick Le Callet
NeurIPS5
2018 Quantifying the Influence of Devices on Quality of Experience for Video Streaming
abstract
The Internet streaming is changing the way of watching videos for people. Traditional quality assessment on the cable/satellite broadcasting system mainly focused on the perceptual quality. Nowadays, this concept has been extended to Quality of Experience (QoE) which considers also the contextual factors, such as the environment, the display devices, etc. In this study, we focus on the influence of devices on QoE. A subjective experiment was conducted by using our proposed AccAnn methodology. The observers evaluated the QoE of the video sequences by considering their Acceptance and Annoyance. Two devices were used in this study, TV and Tablet. The experimental results showed that the device was a significant influence factor on QoE. In addition, we found that this influence varied with the QoE of the video sequences. To quantify this influence, the Eliminated-By-Aspects model was used. The results could be used for the training of a device-neutral objective QoE metric. For video streaming providers, the quantification results of the influence from devices could be used to optimize the selection of streaming content. On one hand it could satisfy the QoE expectations of the observers according to the used devices, on the other hand it could help to save the bitrates.
Jing Li 0026, Lukas Krasula, Patrick Le Callet, Zhi Li 0001, Yoann Baveye
PCS3
2018 Improving the discriminability of standard subjective quality assessment methods: a case study
abstract
Subjective assessment for image or video qualities is considered as the most reliable way to obtain the ground truth for the development of objective quality metrics, especially when leaded by Mean Opinion Score (MOS approaches). However, obtained MOS with standard protocols are noisy due to subject's personal characteristics, such as viewing experience, gender or profession, leading to uncertain ground truth driven by the number of panelists/subjects. The usual way to reduce uncertainty relies on raising this number. In this paper, we demonstrate how a recently introduced Maximum Likelihood Estimation (MLE) based quality recovery model can improve the discriminability of standard subjective quality assessment. Compared to straightforward MOS computation, we present a case study where one can save between 26% to 39% in terms of numbers of subjects at the same discriminability.
Jing Li 0026, Patrick Le Callet
QoMEX2
2018 Introducing UN Salient360! Benchmark: A platform for evaluating visual attention models for 360° contents
abstract
Virtual Reality (VR) provides the users with new immersive media experiences, offering the possibility to freely explore 360° content. Understanding these new exploration behaviors is crucial for the development of efficient techniques for processing, coding, delivering and rendering omnidirectional content to offer the highest possible Quality of Experience (QoE). Progress has already been made on visual attention (VA) modeling for 360° content. In this paper we briefly review the current status of research on this topic that led us to propose a benchmarking platform for evaluating and comparing the performance of models for saliency and scanpath prediction for 360° content. This paper introduces the `UN Salient360! benchmark” platform featuring a dataset, a toolbox and a framework for evaluation of different class of models. This online platform can be found in httns://salient360.ls2n.fr/.
Jesús Gutiérrez 0001, Erwan J. David, Antoine Coutrot, Matthieu Perreira Da Silva, Patrick Le Callet
QoMEX5
2018 Proof-of-concept: role of generic content characteristics in optimizing video encoders - Complexity- and content-aware sequence-level encoder parameter decision framework
Ahmed Aldahdooh, Marcus Barkowsky, Patrick Le Callet
Multim. Tools Appl.3
2018 MPEG-4 AVC stream-based saliency detection. Application to robust watermarking
Marwa Ammar, Mihai Mitrea, Marwen Hasnaoui, Patrick Le Callet
Signal Process. Image Commun.4
2018 Toolbox and dataset for the development of saliency and scanpath models for omnidirectional/360° still images
Jesús Gutiérrez 0001, Erwan J. David, Yashas Rai, Patrick Le Callet
Signal Process. Image Commun.4
2018 Object Shape Approximation and Contour Adaptive Depth Image Coding for Virtual View Synthesis
abstract
A depth image provides partial geometric information of a 3D scene, namely the shapes of physical objects as observed from a particular viewpoint. This information is important when synthesizing images of different virtual camera viewpoints via depth-image-based rendering (DIBR). It has been shown that depth images can be efficiently coded using contour-adaptive codecs that preserve edge sharpness, resulting in visually pleasing DIBR-synthesized images. However, contours are typically losslessly coded as side information, which is expensive if the object shapes are complex. In this paper, we pursue a new paradigm in depth image coding for color-plus-depth representation of a 3D scene: in a pre-processing step, we pro-actively simplify object shapes in a depth and color image pair to reduce depth coding cost, at a penalty of a slight increase in synthesized view distortion. Specifically, we first mathematically derive a distortion upper-bound proxy for 3DSwIM—a quality metric tailored for DIBR-synthesized images. This proxy reduces inter-dependency among pixel rows in a block to ease optimization. We then approximate object contours via a dynamic programming algorithm to optimally tradeoff coding the cost of contours using arithmetic edge coding with our proposed view synthesis distortion proxy. We modify the depth and color images according to the approximated object contours in an inter-view consistent manner. These are then coded, respectively, using a contour-adaptive image codec based on graph Fourier transform for edge preservation and High Efficiency Video Coding (HEVC) intra. Experimental results show that by maintaining sharp but simplified object contours during contour-adaptive coding, for the same visual quality of DIBR-synthesized virtual views, our proposal can reduce depth image coding rate by up to 22% in 3DSwIM and 42% in peak signal-to-noise ratio compared with alternative coding strategies, such as HEVC intra.
Yuan Yuan 0007, Gene Cheung, Patrick Le Callet, Pascal Frossard, H. Vicky Zhao
IEEE Trans. Circuits Syst. Video Technol.3
2018 Data Analysis in Multimedia Quality Assessment: Revisiting the Statistical Tests
abstract
Assessment of multimedia quality relies heavily on subjective assessment, and is typically done by human subjects in the form of preferences or continuous ratings. Such data are crucial for analysis of different multimedia-processing algorithms as well as validation of objective (computational) methods for the said purpose. To that end, statistical testing provides a theoretical framework toward drawing meaningful inferences, and making well-grounded conclusions and recommendations. While parametric tests (such as t test, ANOVA, and error estimates like confidence intervals) are popular and widely used in the community, there appears to be a certain degree of confusion in the application of such tests. Specifically, the assumptions of normality and homogeneity of variance are often not well understood, leading to incorrect application and/or interpretation of the statistical test results. Therefore, the main goal of this paper is to present new guidelines toward proper use of statistical tests and, hence, fix some of the issues in multimedia quality assessment. The said guidelines are derived based on theoretical analysis of sampling distribution of test statistics, and consider practical aspects of data analysis in the said domain. Experimental results on both simulated and real data are presented to support the arguments made. Software that implements the said recommendations is also made publicly available, in order to help researchers and practitioners perform correct statistical comparison of models.
Manish Narwaria, Lukas Krasula, Patrick Le Callet
IEEE Trans. Multim.3
2017 Inpainting-based error concealment for low-delay video communication
abstract
Error concealment (EC) is one of the target applications of inpainting techniques. Some methods combine the estimated lost motion vectors (MVs) with the exemplar-based inpainting technique to recover the lost regions. Due to the erroneous motion vectors that might indicate a moving object as background object and vice versa, these methods are still showing visual artifacts in the recovered regions. In this paper, a concept of motion map that can be easily generated in the decoder side is introduced and it is combined with the exemplar-based inpainting technique. The proposed method introduces an adaptive search window size that trades-off the quality and complexity. Moreover, an optional blending technique is proposed to limit the spatio-temporal artifacts. Experiments show that the proposed method improves the visual quality with 5dB on average relative to the state-of-the-art inpainting-based EC method.
Ahmed Aldahdooh, Marcus Barkowsky, David Bull 0001, Patrick Le Callet
ICASSP4
2017 Reduced-reference quality metric for screen content image
abstract
With the prevalence of digital products like cellphone, tablet and personal computer, the screen content image (SCI) consisting of text, graphic, and natural scene picture becomes a significant media in various communication scenarios. Consequently, we proposed a reduced-reference quality metric dedicated for SCI. The main contribution includes 2 aspects: 1) we innovatively proposed a layer-based segmentation method to divide SCI into text layer and pictorial layer; 2) we designed respective quality metrics dedicated for text and pictorial layers with a novel pooling strategy considering human visual saliency for SCI. Furthermore, exhaustive experimental results indicate that the proposed metric is highly comparative compared with state-of-the-art full-reference quality metrics.
Zhaohui Che, Guangtao Zhai, Ke Gu 0001, Patrick Le Callet
ICIP4
2017 Using multiscale analysis for blind quality assessment of DIBR-synthesized images
abstract
In this paper we propose to blindly evaluate the quality of images synthesized based on a depth image-based rendering (DIBR) procedure. As an important branch in virtual reality (VR), superior DIBR techniques provide free viewpoints in many real applications such as remote surveillance and education, but few efforts have been made to measure the performance of DIBR methods (i.e. the quality of DIBR-synthesized images), especially in the condition of reference unavailable. To this aim, we put forward a new no-reference (NR) image quality assessment (IQA) model via multiscale analysis, dubbed as MSA. The design philosophy of our proposed MSA model is that the DIBR-introduced geometry distortions damage the self-similarity characteristic of natural images and the damage degrees present regular variations at distinct scales. Through systematically incorporating the measurements of the variations provided above, our MSA model can faithfully predict the quality of images generated using different DIBR technologies. Results of experiments demonstrate that the proposed blind MSA model has delivered noticeably better performance than state-of-the-art full-and no-reference IQA methods.
Ke Gu 0001, Junfei Qiao 0001, Patrick Le Callet, Zhifang Xia, Weisi Lin
ICIP3
2017 Video quality enhancement via QP adaptation based on perceptual coding maps
abstract
This paper introduces a method for adapting block quantisation parameter values in HEVC video compression based on perceptual coding maps. These maps are computed per block taking into account masking effects. Masking levels are calculated using spatial, temporal and foveation features that are extracted from the video and are stored in a perceptual coding map. The produced map drives a QP adaptation process that aims to redistribute coding bits in the frame so that the perceived quality is improved, especially at those bitrates where coding artifacts become visible (mid to high QP values). The subjective performance evaluation that was conducted showed that the proposed method can offer a measurable improvement in perceived quality relative to a constant QP approach, with Bjontegaard mean opinion scores (MOS) gains reaching almost 9% for the test sequences used. The paper additionally highlights the need for further work in order to increase gains in perceived quality and optimise parameter selection.
Miltiadis Alexios Papadopoulos, Yashas Rai, Angeliki V. Katsenou, Dimitris Agrafiotis, Patrick Le Callet, David Bull 0001
ICIP5
2017 CVIQD: Subjective quality evaluation of compressed virtual reality images
abstract
The 360-degree spherical images/videos, also called Virtual Reality (VR) images/videos, can provide immersive experience of the real-world scenes in some specific systems. This makes it widely employed in concerts/sports events live and VR movies. However, it is difficult to transport, compress or store VR images/videos due to their high resolution. So it is significant to research how the popular coding technologies influence the quality of VR images. To this aim, this paper carries out subjective quality evaluation of compressed VR images and examines the correlation performance of popular objective quality measures in accordance with the aforesaid subjective ratings. We first establish a Compressed VR Image Quality Database (CVIQD), which includes five source VR images and associated 165 compressed images under three prevailing coding technologies. The Single-Stimulus (SS) method is exploited to collect the subjective scores from 20 inexperienced viewers. Next, we implement 10 classical and recent objective quality metrics on the CVIQD database and compute the correlation between each above quality metric and subjective assessment in terms of five commonly used performance indices. Experimental results reveal that multi-scale based MS-SSIM and ADD-SSIM models have lead to high correlation with human visual perception.
Wei Sun 0029, Ke Gu 0001, Guangtao Zhai, Siwei Ma 0001, Weisi Lin, Patrick Le Callet
ICIP6
2017 Image quality assessment for free viewpoint video based on mid-level contours feature
abstract
Free view point video (FVV), which offers immersive experience to users with multiple views, is one of the new trends in advanced visual media. These new viewpoints are traditionally synthesized via depth image-based rendering(DIBR) and geometric distortions are therefore observed. Mid-level contours descriptors are capable of evaluating such edges incoherence among the synthesized images which common image quality metrics fail to capture. In this paper, we use the concept of ‘Sketch Token’, that is a mid-level contours descriptor, and introduce a novel metric for DIBR-synthesized image quality assessment by measuring how classes of contours change after synthesis. Experiments are conducted on the IRCCyN/IVC DIBR image database and the results show that the proposed metric achieves a correlation of 88.77% which is comparable to state-of-the-art metrics like MW-PSNR and MP-PSNR.
Suiyi Ling, Patrick Le Callet
ICME2
2017 Image Quality Assessment for DIBR Synthesized Views using Elastic Metric
abstract
Frames of free viewpoint video (FVV) synthesized with depth image-based rendering (DIBR) mainly contains special local artifacts like geometric distortions, in which the shape of objects may be stretched/bent. Human observers tend to perceive such local severe deformations instead of consistent shifting artifacts that penalized by most of the existing metrics. Elastic metric is capable of measuring the difference in stretching or bending between two curves, and thus is suitable for evaluating such geometric distortions. In this paper, an elastic metric based image quality assessment (EM-IQA) scheme is proposed by first selecting local distortion regions and then quantifying the deformations of curves. According to the experimental results on the IRCCyn/IVC DIBR image database, the proposed EM-IQA outperforms the state of the art metrics designed for synthesized images and obtains a gain of 6.97% in pearson correlation compared to the second best performing MP-PSNRreduced.
Suiyi Ling, Patrick Le Callet
ACM Multimedia2
2017 Video quality perception in telesurgery
abstract
Telesurgery enables an expert surgeon to assist a remote surgeon during a surgical intervention, which benefits patient care in resource-poor settings. In reality, videos of surgical procedures are compressed and transmitted over large distances in real time and, therefore, are subject to a wide variety of distortions. These distortions degrade the quality of videos and potentially affect the performance of the surgeons. Very little work has been carried out on human perception of video quality in the context of telesurgery. In this paper, we investigate the impact of video compression on the perceived quality of surgical videos. We designed and performed a psychophysical experiment where surgeons rated the quality of surgical videos distorted with two different compression schemes at various compression ratios. Experimental results demonstrate that the impact of video content and compression strategy on the perceived quality is statistically significant.
Lucie Lévêque, Hantao Liu, Christine Cavaro-Ménard, Yongqiang Cheng 0001, Patrick Le Callet
MMSP5
2017 Quality assessment for synthesized view based on variable-length context tree
abstract
In free viewpoint television (FTV) application scenario, views that synthesized with depth image-based rendering (DIBR) techniques mainly contain special artifacts like geometric distortions. These artifacts may affect the structure of images/videos by changing the global contour characteristics and thus are annoying for human observers. Context tree based contour coding scheme can be a good tool to measure such structure loss since the more geometric distortion there is in the synthesized view the lager the gap between the encoding cost of the reference and synthesized views. In this paper, we investigate whether such overall encoding cost can be related to perceptual annoyance reflecting in quality score and propose a variable-length context tree based image quality assessment (CT-IQA) scheme. This scheme quantify (1) the overall structure dissimilarity and (2) dissimilarities in various contour characteristics between the reference and synthesized views. The proposed metric is robust to global shifting artifact that is over penalized by traditional metrics. According to the experimental results on the IRCCyn/IVC DIBR image database, the performance of the proposed CT-IQA is promising.
Suiyi Ling, Patrick Le Callet, Gene Cheung
MMSP2
2017 A Dataset of Head and Eye Movements for 360 Degree Images
abstract
Understanding how observers watch visual stimuli like Images and Videos has helped the multimedia encoding, transmission, quality assessment and rendering communities immensely, to learn the regions important to an observer and provide to him/her an optimum quality of experience. The problem is even more paramount in case of 360 degree stimuli considering that most/a part of the content might not be seen by the observers at all, while other regions maybe extraordinarily important. Attention studies in this area has however been missing, mainly due to the lack of a dataset and guidelines to evaluate and compare visual attention/saliency in such scenarios. In this work, we present a dataset of sixty different 360 degree images, each watched by at-least 40 observers. Additionally, we also provide guidelines and tools to the community regarding the procedure to evaluate and compare saliency in omni-directional images. Some basic image/ observer agnostic viewing characteristics, like variation of exploration strategies with time and expertise, and also the effect of eye-movement within the view-port are explored. The dataset and tools are made available for free use by the community and is expected to promote Reproducible Research for all future work on computational modeling of attention in 360 scenarios.
Yashas Rai, Jesús Gutiérrez 0001, Patrick Le Callet
MMSys3
2017 Characterization and selection of light field content for perceptual assessment
abstract
Light field technology may have a positive impact on several multimedia applications thanks to novel ways to explore the captured scenes, such as changing the parallax (horizontally and vertically) and refocusing the content. These innovative use cases require new considerations that affect the whole processing chain, from content acquisition to visualization, as well as the methodologies for quality evaluation. In particular, capturing and selecting the appropriate content is crucial for a successful development and evaluation of audiovisual technologies. Thus, this paper presents a framework addressing the reconsideration of the space of attributes for an adequate characterization of light field data. Firstly, an exhaustive characterization of light field content is described, based on various particular features, including depth and refocusing properties to traditional spatial, temporal, and color indicators. Then, based on this characterization, specific techniques are proposed for an effective selection of light field content for perceptual quality assessment.
Pradip Paudyal, Jesús Gutiérrez 0001, Patrick Le Callet, Marco Carli, Federica Battisti
QoMEX3
2017 Which saliency weighting for omni directional image quality assessment?
abstract
With the explosion of Virtual Reality technologies, the production and usage of omni directional images (a.k.a 360 images) is presenting new challenges in the domains of compression, transmission and rendering. The evaluation of the quality of images generated by these technologies is therefore paramount. As the exploration of 360 images within a Head-Mounted Display (HMD) is non-uniform, current state of the art proposes a saliency weighting of distortions (between a reference and an impaired version) thus allowing us to highlight impairments in frequently attended regions. So far, saliency maps have been generated by tracking head motion alone, and consider that the view-port orientation alone is sufficient to determine saliency. The added value of eye gaze tracking within the viewport has not been studied in this domain yet. In this work, an eye-tracking experiment is performed using a HMD and is followed by subsequent gaze analysis to appreciate the visual attention behavior within a view-port. Results suggest that most eye-gaze fixations are rather far away from the center of the viewport. Across contents and observers, gaze fixations are quasi-isotropically distributed in orientation. The average location of gaze fixation (across contents and observers) from the center of the view-port varies between 14 and 20 visual degrees and these values correspond to a shift in retina beyond the para and peri fovea, into the extra-perifoveal region. A saliency weighting model based on foveation, centered at the middle of the view-port seems to be a correct assumption in only 2.5% of the overall scenarios and is consequently questionable. Therefore there is a need to refine saliency modeling and weighting for quality assessment in case of panoramic viewing.
Yashas Rai, Patrick Le Callet, Philippe Guillotel
QoMEX2
2017 Learning visual saliency from human fixations for stereoscopic images
Yuming Fang 0001, Jianjun Lei 0001, Jia Li 0003, Long Xu 0001, Weisi Lin, Patrick Le Callet
Neurocomputing6
2017 Visual Attention Modeling for Stereoscopic Video: A Benchmark and Computational Model
abstract
In this paper, we investigate the visual attention modeling for stereoscopic video from the following two aspects. First, we build one large-scale eye tracking database as the benchmark of visual attention modeling for stereoscopic video. The database includes 47 video sequences and their corresponding eye fixation data. Second, we propose a novel computational model of visual attention for stereoscopic video based on Gestalt theory. In the proposed model, we extract the low-level features, including luminance, color, texture, and depth, from discrete cosine transform coefficients, which are used to calculate feature contrast for the spatial saliency computation. The temporal saliency is calculated by the motion contrast from the planar and depth motion features in the stereoscopic video sequences. The final saliency is estimated by fusing the spatial and temporal saliency with uncertainty weighting, which is estimated by the laws of proximity, continuity, and common fate in Gestalt theory. Experimental results show that the proposed method outperforms the state-of-the-art stereoscopic video saliency detection models on our built large-scale eye tracking database and one other database (DML-ITRACK-3D).
Yuming Fang 0001, Chi Zhang 0027, Jing Li 0026, Jianjun Lei 0001, Matthieu Perreira Da Silva, Patrick Le Callet
IEEE Trans. Image Process.6
2017 Quality Assessment of Sharpened Images: Challenges, Methodology, and Objective Metrics
abstract
Most of the effort in image quality assessment (QA) has been so far dedicated to the degradation of the image. However, there are also many algorithms in the image processing chain that can enhance the quality of an input image. These include procedures for contrast enhancement, deblurring, sharpening, up-sampling, denoising, transfer function compensation, etc. In this work, possible strategies for the quality assessment of sharpened images are investigated. This task is not trivial because the sharpening techniques can increase the perceived quality, as well as introduce artifacts leading to the quality drop (over-sharpening). Here, the framework specifically adapted for the quality assessment of sharpened images and objective metrics comparison in this context is introduced. However, the framework can be adopted in other quality assessment areas as well. The problem of selecting the correct procedure for subjective evaluation was addressed and a subjective test on blurred, sharpened, and over-sharpened images was performed in order to demonstrate the use of the framework. The obtained ground-truth data were used for testing the suitability of state-ofthe- art objective quality metrics for the assessment of sharpened images. The comparison was performed by novel procedure using ROC analyses which is found more appropriate for the task than standard methods. Furthermore, seven possible augmentations of the no-reference S3 metric adapted for sharpened images are proposed. The performance of the metric is significantly improved and also superior over the rest of the tested quality criteria with respect to the subjective data.
Lukas Krasula, Patrick Le Callet, Karel Fliegel, Milos Klima
IEEE Trans. Image Process.2
2016 Transform Coding for On-the-Fly Learning Based Block Transforms
abstract
Block based video codecs developed in last decades have employed Discrete Cosine Transform (DCT) as a candidate transform for compacting the energy of the predicted residual data into fewer coefficients. However, DCT is not optimal for blocks with directional structure and high texture. Content-adaptive transforms can be used to achieve better compression. However, these adaptive transforms must be sent to the decoder, increasing the bit-rate in return. It is therefore important to take care of the overhead of sending the transforms in order to benefit from the better compaction of the adaptive transforms. This paper addresses this issue and provides methods to efficiently encode these adaptive transforms that are learned from the content. A statistical model is derived to determine the precision required to encode a transform with a negligible loss in the performance of adaptive transforms. Results shows that the overhead can be reduced by an average 70% for low QP and 86% for high QP values compared to encoding the complete transform with a little performance loss.
Saurabh Puri, Sebastien Lasserre, Patrick Le Callet, Fabrice Le Léannec
DCC3
2016 Annealed learning based block transforms for HEVC video coding
abstract
Most of the recent video compression standards employ the Discrete Cosine Transform (DCT) for transforming the residual signal in order to remove spatial correlation and to achieve higher compression efficiency. However, by careful adaptation of transforms to the video content, a better set of integer transforms can be obtained. This paper proposes a new on-the-fly block-based transform optimization technique which involves first the classification of the residual blocks based on the cost of encoding the block, and then the generation of new optimized transforms for each class. An annealing based learning technique is further proposed in this paper in order to improve the performance of the optimization algorithm. The algorithm is tested using the latest HEVC test software where an optimized set of transforms is learned on the first frame of the HEVC test sequences and then applied to the subsequent frames in a Random Access (RA) and All Intra (AI) configuration. The results shows that this method can gain over 2% in terms of Bjontegaard Delta (BD)-rate compared to standard HEVC encoder in AI configuration and nearly 1.5% in RA.
Saurabh Puri, Sebastien Lasserre, Patrick Le Callet
ICASSP3
2016 Visual saliency detection via image complexity feature
abstract
In this paper we propose a novel bottom-up visual saliency detection model by analysis of image complexity. Compared with existing works, we emphasize the important impact of image complexity on saliency detection. Inspired by the free energy theory, a hybrid parametric and non-parametric model is used to estimate the complexity of a visual signal. Taking the image complexity as a new feature, this paper constructs a heuristic framework to systematically combine two different types of saliency detection models, separately using local and global features, in order to predict human fixation points more accurately. In contrast to classical and modern models, our algorithm has achieved noticeably superior results. And furthermore, it is worthy to stress that the proposed saliency detection method can also help to facilitate the performance of image quality metrics on popular image databases.
Min Liu 0003, Ke Gu 0001, Guangtao Zhai, Patrick Le Callet
ICIP4
2016 Modeling the perceptual distortion of dynamic textures and its application in HEVC
abstract
In the quest of perceptually optimized video coding, coding textures is representing a challenging case. While a large body of research was put into the perception of static textures, dynamic textures are still not sufficiently explored. In this paper, we focus on short term consistent patches, known as dynamic textures, with a very limited spatial and temporal extent. We estimated the visual distortion due to HEVC compression via subjective testing. The estimated distortion profile per texture was used to optimize the HEVC coding process. Experimental results showed that this technique can offer a high bitrate saving for the same subjective quality.
Karam Naser, Vincent Ricordel, Patrick Le Callet
ICIP3
2016 Role of HEVC coding artifacts on gaze prediction in interactive video streaming systems
abstract
Sensitivity to spatial details drops across the visual periphery, and hence video streaming systems that gracefully degrades quality away from the viewpoint of the observer, provides an optimum viewing experience with potentially large bitrate savings. As reaction latency is an important performance parameter of such systems, good prediction of future gaze locations at the transmission end is very important. A major research question here is: whether a gaze prediction model designed using a pristine undistorted video, is also able to predict the gaze pattern of users when they watch a distorted/ adaptively distorted version of the same video. With several improvements to existing gaze prediction schemes, in combination with a controlled subjective experiment, we confirm not only that HEVC coding distortions have no significant impact on the predictability of gaze patterns, but also that gaze prediction errors can be restricted to 1.5 degrees of viewing angle for a round trip delay of up to 200ms.
Yashas Rai, Patrick Le Callet, Gene Cheung
ICIP2
2016 Impact of visual angle on attention deployment and robustness of visual saliency models in videos: From SD to UHD
abstract
The emergence of UHD video format induces larger screens and involves a wider stimulated visual angle. Therefore, its effect on visual attention can be questioned since it can impact quality assessment, metrics but also the whole chain of video processing and creation. Moreover, changes in visual attention from different viewing conditions challenge visual attention models. In this paper, we present a comparative study of visual attention and viewing behavior on three video datasets in SD, HD and UHD conditions. Then, we propose and assess an improvement for video visual attention models by applying a stimulated visual angle dependent center model.
Toinon Vigier, Matthieu Perreira Da Silva, Patrick Le Callet
ICIP3
2016 Deep Learning for Image Memorability Prediction: the Emotional Bias
abstract
Image memorability prediction is a recent topic in computer science. First attempts have shown that it is possible to computationally infer from the intrinsic properties of an image the extent to which it is memorable. In this paper, we introduce a fine-tuned deep learning-based computational model for image memorability prediction. The performance of this model significantly outperforms previous work and obtains a 32.78% relative increase compared to the best-performing model from the state of the art on the same dataset. We also investigate how our model generalizes on a new dataset of 150 images, for which memorability and affective scores were collected from 50 participants. The prediction performance is weaker on this new dataset, which highlights the issue of representativity of the datasets. In particular, the model obtains a higher predictive performance for arousing negative pictures than for neutral or arousing positive ones, recalling how important it is for a memorability dataset to consist of images that are appropriately distributed within the emotional space.
Yoann Baveye, Romain Cohendet, Matthieu Perreira Da Silva, Patrick Le Callet
ACM Multimedia4
2016 Effect of content features on short-term video quality in the visual periphery
abstract
The area outside our central field of vision, also referred to as the visual periphery, captures most information in a visual scene, although much less sensitive than the central Fovea. Vision studies in the past have stated that there is reduced sensitivity of texture, color, motion and flicker (temporal harmonic) perception in this area, that bears an interesting application in the domain of quality perception. In this work, we particularly analyze the perceived subjective quality of videos containing H.264/AVC transmission impairments, incident at various degrees of retinal eccentricities of observers. We relate the perceived drop in quality, to five basic types of features that are important from a perceptive standpoint: texture, color, flicker, motion trajectory distortions and also the semantic importance of the underlying regions. We are able to observe that the perceived drop in quality across the visual periphery, is closely related to the Cortical Magnification fall-off characteristics of the V1 cortical region. Additionally, we see that while object importance and low frequency spatial distortions are important indicators of quality in the central foveal region, temporal flicker and color distortions are the most important determinants of quality in the periphery. We therefore conclude that, although users are more forgiving of distortions they viewed peripherally, they are nevertheless not totally blind towards it: the effects of flicker and color distortions being particularly important.
Yashas Rai, Ahmed Aldahdooh, Suiyi Ling, Marcus Barkowsky, Patrick Le Callet
MMSP5
2016 A new HD and UHD video eye tracking dataset
abstract
The emergence of UHD video format induces larger screens and involves a wider stimulated visual angle. Therefore, its effect on visual attention can be questioned since it can impact quality assessment, metrics but also the whole chain of video processing and creation. Moreover, changes in visual attention from different viewing conditions challenge visual attention models. In this paper, we present a new HD and UHD video eye tracking dataset composed of 37 high quality videos observed by more than 35 naive observers. This dataset can be used to compare viewing behavior and visual saliency in HD and UHD, as well as for any study on dynamic visual attention in videos. It is available at http://ivc.univ-nantes.fr/en/databases/HD_UHD_Eyetracking_Videos/.
Toinon Vigier, Josselin Rousseau, Matthieu Perreira Da Silva, Patrick Le Callet
MMSys4
2016 A foveated short term distortion model for perceptually optimized dynamic textures compression in HEVC
abstract
Due to their rapid change over time, dynamic textures represent challenging contents for the video compression standards. While several studies proposed different coding techniques, such as synthesis based coding, they mostly lack the perceptual studies on the perceived distortions. In this paper, a framework of perceptually optimized dynamic texture compression is presented. It is based on a psychophysically driven distortion model, which is utilized inside the rate distortion loop of HEVC. The model is tested for three compression levels, and showed a significant rate saving for both training and validation sequences set.
Karam Naser, Vincent Ricordel, Patrick Le Callet
PCS3
2016 Improved coefficient coding for adaptive transforms in HEVC
abstract
Adaptive transform learning schemes have been extensively studied in the literature with a goal to achieve better compression efficiency compared to extensively used Discrete Cosine Transforms (DCT) inside a video codec. These transforms are learned offline on a large training set and are tested either in competition with or in place of the core transforms i.e. DCT. In our previous work, we proposed an alternative approach where a set of non-separable content-adaptive transforms are learned on-the-fly on a sequence and then tested on the same sequence in competition with the core transforms. In this paper, we propose to further improve the previously proposed learning scheme by improving the coding of transformed coefficients in context to the online learning of transforms. The first proposed method improves the convergence of the learning scheme by re-ordering the transformed coefficients at each iteration. The second proposed method improves the compression efficiency by modifying the coding of last significant coefficient position in context to adaptive transforms. The results shows that by combining the above proposed methods, one can achieve 1.2% gain on top of the previously proposed scheme.
Saurabh Puri, Sebastien Lasserre, Patrick Le Callet
PCS3
2016 Spatio-temporal error concealment technique for high order multiple description coding schemes including subjective assessment
abstract
Error resilience (ER) is an important tool in video coding to maximize the quality of Experience (QoE). The prediction process in video coding became complex which yields an unsatisfying video quality when NALunit packets are lost in error-prone channels. There are different ER techniques and multiple description coding (MDC) is one of the promising technique for this problem. MDC is categorized into different types and, in this paper, we focus on temporal MDC techniques. In this paper, a new temporal MDC scheme is proposed. In the encoding process, the encoded descriptions contain primary frames and secondary frames (redundant representations). The secondary frames represent the MVs that are predicted from previous primary frames such that the residual signal is set to zero and is not part of the rate distortion optimization. In the decoding process of the lost frames, a weighted average error concealment (EC) strategy is proposed to conceal these frames. The proposed scheme is subjectively evaluated along with other schemes and the results show that the proposed scheme is significantly different from most of other temporal MDC schemes.
Ahmed Aldahdooh, Marcus Barkowsky, Patrick Le Callet
QoMEX3
2016 Using individual data to characterize emotional user experience and its memorability: Focus on gender factor
abstract
Delivering the same digital image to several users is not necessarily providing them the same experience. In this study, we focused on how different affective experiences impact the memorability of an image. Forty-nine participants took part in an experiment in which they saw a stream of images conveying various emotions. One day later, they had to recognize the images displayed the day before and rate them according to the positivity/ negativity of the emotional experience the images induced. In order to better appreciate the underlying idiosyncratic factors that affect the experience under test, prior to the test session we collected not only personal information but also results of psychological tests to characterize individuals according to their dominant personality in terms of masculinity-femininity (Bem Sex Role Inventory) and to measure their emotional state. The results show that the way an emotional experience is rated depends on personality rather than biological sex, suggesting that personality could be a mediator in the well-established differences in how males and females experience emotional material. From the collected data, we derive a model including individual factors relevant to characterize the memorability of the images, in particular through the emotional experience they induced.
Romain Cohendet, Anne-Laure Gilet, Matthieu Perreira Da Silva, Patrick Le Callet
QoMEX4
2016 How to benchmark objective quality metrics from paired comparison data?
abstract
The procedures commonly used to evaluate the performance of objective quality metrics rely on ground truth mean opinion scores and associated confidence intervals, which are usually obtained via direct scaling methods. However, indirect scaling methods, such as the paired comparison method, can also be used to collect ground truth preference scores. Indirect scaling methods have a higher discriminatory power and are gaining popularity, for example in crowdsourcing evaluations. In this paper, we present how the classification errors, an existing analysis tool, can also be used with subjective preference scores. Additionally, we propose a new analysis tool based on the receiver operating characteristic analysis. This tool can be used to further assess the performance of objective metrics based on ground truth preference scores. We provide a MATLAB script with an implementation of the proposed tools and we show one example of application of the proposed tools.
Philippe Hanhart, Lukas Krasula, Patrick Le Callet, Touradj Ebrahimi
QoMEX3
2016 On the accuracy of objective image and video quality models: New methodology for performance evaluation
abstract
There are several standard methods for evaluating the performance of models for objective quality assessment with respect to results of subjective tests. However, all of them suffer from one or more of the following drawbacks: They do not consider the uncertainty in the subjective scores, requiring the models to make certain decision where the correct behavior is not known. They are vulnerable to the quality range of the stimuli in the experiments. In order to compare the models, they require a mapping of predicted values to the subjective scores, thus not comparing the models exactly as they are used in the real scenarios. In this paper, new methodology for objective models performance evaluation is proposed. The method is based on determining the classification abilities of the models considering two scenarios inspired by the real applications. It does not suffer from the previously stated drawbacks and enables to easily evaluate the performance on the data from multiple subjective experiments. Moreover, techniques to determine statistical significance of the performance differences are suggested. The proposed framework is tested on several selected metrics and datasets, showing the ability to provide a complementary information about the models' behavior while being in parallel with other state-of-the-art methods.
Lukas Krasula, Karel Fliegel, Patrick Le Callet, Milos Klima
QoMEX3
2016 Estimation of perceptual redundancies of HEVC encoded dynamic textures
abstract
Statistical redundancies have been the dominant target in the image/video compression standards. Perceptually, there exists further redundancies that can be removed to further enhance the compression efficiency. In this paper, we considered short term homogeneous patches that fall into the foveal vision as dynamic textures, for which a psychophysical test was used to estimate their amount of perceptual redundancies. We demonstrated the possible rate saving by utilizing these redundancies. We further designed a learning model that can precisely predict the amount of redundancies and accordingly proposed a generalized perceptual optimization framework.
Karam Naser, Vincent Ricordel, Patrick Le Callet
QoMEX3
2016 Free viewpoint video quality assessment based on morphological multiscale metrics
abstract
In this paper, two image quality metrics based on morphological multiscale decompositions have been applied for the evaluation of free viewpoint video sequences. These sequences are synthesized using decompressed depth maps in the Depth-Image-Based Rendering synthesis process. Since the synthesis introduces edge distortion in the synthesized image/video, morphological filters are used for their ability to maintain important geometric information (i.e., edges) across different resolution levels. The edges displacement in different resolution scales is evaluated by means of the Mean Square Error. The adopted metrics show higher correlation with human judgment than state-of-the-art image quality measures used in this context.
Dragana Sandic-Stankovic, Federica Battisti, Dragan Kukolj, Patrick Le Callet, Marco Carli
QoMEX4
2016 Visual attention as a dimension of QoE: Subtitles in UHD videos
abstract
With the ever-growing availability of multimedia content produced, broadcast and consumed worldwide, subtitling is becoming an essential service to quickly share understandable content. Simultaneously, the increased resolution of the ultra high definition (UHD) standard comes with wider screens and new viewing conditions. Services as the display of subtitles thus require adaptation to better fit the new induced viewing visual angle. This paper aims at evaluating quality of experience of subtitled movies in UHD to propose guidelines for the appearance of subtitles. From an eye-tracking experiment conducted on 68 observers and 30 video sequences, viewing behavior and visual saliency are analyzed with and without subtitles and for different subtitle styles. Various metrics based on eye-tracking data, such as the Reading Index for Dynamic Texts (RIDT), are computed to objectively measure the ease of reading and subtitle disturbance. The results mainly show that doubling the visual angle of subtitles from HD to UHD guarantees subtitle readability without compromising the enjoyment of the video content.
Toinon Vigier, Yoann Baveye, Josselin Rousseau, Patrick Le Callet
QoMEX4
2016 Computational modeling of artistic intention: Quantify lighting surprise for painting analysis
abstract
The use of strong lighting contrast to accentuate objects and figures in a painting—called Chiaroscuro—is popular among Renaissance painters such as Caravaggio, La Tour and Rembrandt. In this paper, we propose a new metric called LuCo to quantify the extent to which Chiaroscuro is employed by an artist in a painting. This measurement could be used to assess the capability of any system to fulfill the original artistic intention and consequently ensure minimal disruptions of Quality of Experience. We first argue that Chiaroscuro is a device for artists to draw attention to specific spatial regions; thus it can be understood as a restricted notion of visual saliency computed using only luminance features. Operationally, using a set of local luminance patches we first compute a Bayesian surprise value, where the prior and posterior probabilities are computed assuming a Gaussian Markov Random Field (GMRF) model. Inverse covariance matrices of the GMRF model are estimated via sparse graph learning for robustness. We construct a histogram using the computed surprise values from different local patches in a painting. Finally, we compute a skewness parameter for the constructed histogram as our LuCo score: large skewness means luminance surprises are either very small or very large, meaning that the artist accentuated lighting contrast in the painting. Experimental results show that paintings by Chiaroscuro artists have higher LuCo scores than 19th century French Impressionists, and Rembrandt's self-portraits have increasingly higher LuCo scores as he aged except for his late period—both trends are in agreement with art historians' interpretations.
Saboya Yang, Gene Cheung, Patrick Le Callet, Jiaying Liu 0001, Zongming Guo
QoMEX3
2016 Saliency-based stereoscopic image retargeting
Yuming Fang 0001, Junle Wang, Yuan Yuan 0029, Jianjun Lei 0001, Weisi Lin, Patrick Le Callet
Inf. Sci.6
2016 A Universal Framework for Salient Object Detection
abstract
In this paper, we propose a novel universal framework for salient object detection, which aims to enhance the performance of any existing saliency detection method. First, rough salient regions are extracted from any existing saliency detection model with distance weighting, adaptive binarization, and morphological closing. With the superpixel segmentation, a Bayesian decision model is adopted to refine the rough saliency map to obtain a more accurate saliency map. An iterative optimization method is designed to obtain better saliency results by exploiting the characteristics of the output saliency map each time. Through the iterative optimization process, the rough saliency map is updated step by step with better and better performance until an optimal saliency map is obtained. Experimental results on the public salient object detection datasets with ground truth demonstrate the promising performance of the proposed universal framework subjectively and objectively.
Jianjun Lei 0001, Bingren Wang, Yuming Fang 0001, Weisi Lin, Patrick Le Callet, Nam Ling, Chunping Hou
IEEE Trans. Multim.5
2016 The Application of Visual Saliency Models in Objective Image Quality Assessment: A Statistical Evaluation
abstract
Advances in image quality assessment have shown the potential added value of including visual attention aspects in its objective assessment. Numerous models of visual saliency are implemented and integrated in different image quality metrics (IQMs), but the gain in reliability of the resulting IQMs varies to a large extent. The causes and the trends of this variation would be highly beneficial for further improvement of IQMs, but are not fully understood. In this paper, an exhaustive statistical evaluation is conducted to justify the added value of computational saliency in objective image quality assessment, using 20 state-of-the-art saliency models and 12 best-known IQMs. Quantitative results show that the difference in predicting human fixations between saliency models is sufficient to yield a significant difference in performance gain when adding these saliency models to IQMs. However, surprisingly, the extent to which an IQM can profit from adding a saliency model does not appear to have direct relevance to how well this saliency model can predict human fixations. Our statistical analysis provides useful guidance for applying saliency models in IQMs, in terms of the effect of saliency model dependence, IQM dependence, and image distortion dependence. The testbed and software are made publicly available to the research community.
Wei Zhang 0072, Ali Borji, Zhou Wang 0001, Patrick Le Callet, Hantao Liu
IEEE Trans. Neural Networks Learn. Syst.4
2015 A multi-slice model observer for medical image quality assessment
abstract
Model observers (MOs) have been developed for the medical image quality assessment. Nowadays, numerous modern medical instruments are capable of producing 3D images, while few researchers have conducted MO studies on 3D data. In this paper, we propose a multi-slice MO when considering a relatively more realistic diagnostic task: the detection-localization of simulated multiple-sclerosis (MS) lesions on 3D magnetic resonance (MR) images. The jackknife free-response receiver operating characteristic (JAFROC) method was used to quantitatively analyse its performances and compare them with those of human observers. Our preliminary results showed that the proposed framework has the potential to approach human detection-localization task performance.
Lu Zhang 0037, Christine Cavaro-Ménard, Patrick Le Callet, Di Ge
ICASSP3
2015 UHD image reconstruction by estimating interpolation error
abstract
With the emerging ultra high definition (UHD) TV technology, the reconstruction of UHD images from transmitted HD images on the receiver's side is of great interests to save bandwidth and be compatible to existing HDTV systems. In this paper, a UHD image reconstruction algorithm is proposed for the UHDTV broadcasting system. A transmitted HD image at the receiver side is first up sampled to UHD resolution, then the error map between the interpolated UHD image and the original UHD image is estimated. The final reconstructed UHD image is the sum of the interpolated UHD image and the estimated UHD error map. There are two steps to estimate the content of the UHD error map: 1) key point detection and prediction estimate the location of the pixels with large error and create a spline tube between two adjacent key points; 2) spline-tube interpolation estimate the error value along the spline tube. Our simulation results on six images show that the proposed reconstruction method performs better than conventional up sampling without estimating the error map.
Kai Berger, Kongfeng Berger, Patrick Le Callet
ICIP3
2015 Contour approximation & depth image coding for virtual view synthesis
abstract
A depth image provides geometric information of a 3D scene, namely the shapes of physical objects captured from a particular viewpoint. This information is important for synthesizing images corresponding to different virtual camera viewpoints via depth-image-based rendering (DIBR). Since it has been shown that blurring of object contours in the depth images leads to bleeding artefacts in virtual images. The most effective way to compress depth images relies on edge-adaptive image codecs that preserve contours, which are losslessly coded as side information (SI). However, lossless coding of the exact object contours can be expensive. In this paper, we argue that the contours themselves can be suitably approximated to save bits, while the depth images piecewise smooth (PWS) characteristic stays preserved. Specifically, we first propose a metric that estimates contour coding rate based on edge statistics. Given an initial rate estimate, we then pro-actively approximate object contours in a way that guarantees rate reduction when coded using arithmetic edge coding (AEC) as SI. Given the sharp but approximated contours, we finally encode the image using an edge-adaptive image codec with graph Fourier transform (GFT) for edge preservation. We show in our experiments that by maintaining sharp but slightly inaccurate object contours, the resulting quality of virtual views synthesized via DIBR exceeds those synthesized using depth images compressed with edge-adaptive codecs that losslessly encode object contours as SI, in particular when the total coding rate budget is low. This confirms that optimized coding of depth images results in an effective tradeoff in the representation of contour and respective depth information.
Yuan Yuan 0007, Gene Cheung, Pascal Frossard, Patrick Le Callet, H. Vicky Zhao
MMSP4
2015 Directional regularity for visual quality estimation
Delei Liu, Yong Xu 0007, Yuhui Quan, Zhiwen Yu 0002, Patrick Le Callet
Signal Process.5
2015 Objective image quality assessment of 3D synthesized views
Federica Battisti, Emilie Bosc, Marco Carli, Patrick Le Callet, Simone Perugia
Signal Process. Image Commun.4
2015 Perceived interest and overt visual attention in natural images
Ulrich Engelke, Patrick Le Callet
Signal Process. Image Commun.2
2015 HDR-VQM: An objective quality measure for high dynamic range video
Manish Narwaria, Matthieu Perreira Da Silva, Patrick Le Callet
Signal Process. Image Commun.3
2014 Stereoscopic image retargeting based on 3D saliency detection
abstract
In this paper, we propose a novel stereoscopic image retargeting algorithm based on 3D visual saliency detection. A new 3D visual attention model is designed based on 2D visual feature detection, depth feature detection and the modeling of various viewing bias in stereo vision. A geometrically consistent seam carving technique is adopted for retargeting stereo image pair. Experimental results demonstrated that both the proposed visual attention model and the proposed retargeting method outperform the state-of-the-art studies.
Junle Wang, Yuming Fang 0001, Manish Narwaria, Weisi Lin, Patrick Le Callet
ICASSP5
2014 Exploring the effects of 3D visual discomfort on viewers' emotions
abstract
Much of the research on stereoscopic 3D (S3D) QoE has dealt with the effects of image features over user's visual discomfort. Nevertheless, the relationships between visual discomfort and other high-level perceptual dimensions such as the viewers' emotional reactions have not been yet explored. Since emotions play a key role on media reception processes, and especially in entertainment contents consumption, this question deserves to be thoroughly addressed to improve the design of image processing techniques. This paper raises and investigates the possible relationship between image distortions and its impact on the emotion context of S3D. We implemented an experimental design in which participants watched a series of 3D contents with different levels of visual distortions supposed to induce visual discomfort, while self-reported and psychophysiological measures of emotions were recorded. Results showed that physiological correlates of emotions were affected by visual discomfort conditions, and that these measures were more sensitive than traditional self-reported measures.
Miguel Barreda-Ángeles, Romuald Pépion, Emilie Bosc, Patrick Le Callet, Alexandre Pereda-Baños
ICIP4
2014 Optimizing feature pooling and prediction models of VQA algorithms
abstract
In this paper, we propose a strategy to optimize feature pooling and prediction models of video quality assessment (VQA) algorithms with a much smaller number of parameters than methods based on machine learning, such as neural networks. Based on optimization, the proposed mapping strategy is composed of a global linear model for pooling extracted features, a simple linear model for local alignment in which local factors depend on source videos, and a non-linear model for quality calibration. Also, a reduced-reference VQA algorithm is proposed to predict the local factors from the source video. In the IRCCyN/IVC video database of content influence and the LIVE mobile video database, the performance of VQA algorithms is improved significantly by local alignment. The proposed mapping strategy with prediction of local factors outperforms one no-reference VQA metric and is comparable to one full-reference VQA metric. Thus predicting the local factors in local alignment based on video content will be a promising new approach for VQA.
Kongfeng Zhu, Marcus Barkowsky, Minmin Shen, Patrick Le Callet, Dietmar Saupe
ICIP4
2014 Free-viewpoint video sequences: A new challenge for objective quality metrics
abstract
Free-viewpoint television is expected to create a more natural and interactive viewing experience by providing the ability to interactively change the viewpoint to enjoy a 3D scene. To render new virtual viewpoints, free-viewpoint systems rely on view synthesis. However, it is known that most objective metrics fail at predicting perceived quality of synthesized views. Therefore, it is legitimate to question the reliability of commonly used objective metrics to assess the quality of free-viewpoint video (FVV) sequences. In this paper, we analyze the performance of several commonly used objective quality metrics on FVV sequences, which were synthesized from decompressed depth data, using subjective scores as ground truth. Statistical analyses showed that commonly used metrics were not reliable predictors of perceived image quality when different contents and distortions were considered. However, the correlation improved when considering individual conditions, which indicates that the artifacts produced by some view synthesis algorithms might not be correctly handled by current metrics.
Philippe Hanhart, Emilie Bosc, Patrick Le Callet, Touradj Ebrahimi
MMSP3
2014 Robust drift-free bit-rate preserving H.264 watermarking
Wei Chen 0054, Zafar Shahid, Thomas Stütz, Florent Autrusseau, Patrick Le Callet
Multim. Syst.5
2014 Reduced reference image quality assessment using regularity of phase congruency
Delei Liu, Yong Xu 0007, Yuhui Quan, Patrick Le Callet
Signal Process. Image Commun.4
2014 Special issue on advances in high dynamic range video research
Marta Mrak, Robert Cohen, Patrick Le Callet
Signal Process. Image Commun.3
2014 Tone mapping based HDR compression: Does it affect visual experience?
Manish Narwaria, Matthieu Perreira Da Silva, Patrick Le Callet, Romuald Pépion
Signal Process. Image Commun.3
2014 Saliency Detection for Stereoscopic Images
abstract
Many saliency detection models for 2D images have been proposed for various multimedia processing applications during the past decades. Currently, the emerging applications of stereoscopic display require new saliency detection models for salient region extraction. Different from saliency detection for 2D images, the depth feature has to be taken into account in saliency detection for stereoscopic images. In this paper, we propose a novel stereoscopic saliency detection framework based on the feature contrast of color, luminance, texture, and depth. Four types of features, namely color, luminance, texture, and depth, are extracted from discrete cosine transform coefficients for feature contrast calculation. A Gaussian model of the spatial distance between image patches is adopted for consideration of local and global contrast calculation. Then, a new fusion method is designed to combine the feature maps to obtain the final saliency map for stereoscopic images. In addition, we adopt the center bias factor and human visual acuity, the important characteristics of the human visual system, to enhance the final saliency map for stereoscopic images. Experimental results on eye tracking databases show the superior performance of the proposed model over other existing methods.
Yuming Fang 0001, Junle Wang, Manish Narwaria, Patrick Le Callet, Weisi Lin
IEEE Trans. Image Process.4
2013 Adaptive contrast adjustment for postprocessing of tone mapped high dynamic range images
abstract
Tone mapping operators (TMOs) employed to visualize high dynamic range (HDR) content on conventional low dynamic range (LDR) devices suffer from two major drawbacks. First, none of them can faithfully reproduce all the contrast present in HDR images. Second, most of them require one or more parameters which are mostly content specific and their optimal values can be set only via subjective testing. To address these issues, this paper proposes that ‘quality driven’ adaptive contrast enhancement is a practical solution. This is achieved by enhancing the contrast adaptively based on the loss of contrast between the HDR and tone mapped image. Experimental results confirm that the proposed adaptive solution always improves upon the contrast achieved from whatever given TMO parameter settings in the tested images. So it helps to achieve the results of a more optimal TMO parameter setting without the human input.
Manish Narwaria, Matthieu Perreira Da Silva, Patrick Le Callet, Romuald Pépion
ISCAS3
2013 Saliency detection for stereoscopic images
abstract
Saliency detection techniques have been widely used in various 2D multimedia processing applications. Currently, the emerging applications of stereoscopic display require new saliency detection models for stereoscopic images. Different from saliency detection for 2D images, depth features have to be taken into account in saliency detection for stereoscopic images. In this paper, we propose a new stereoscopic saliency detection framework based on the feature contrast of color, intensity, texture, and depth. Four types of features including color, luminance, texture, and depth are extracted from DC-T coefficients to represent the energy for image patches. A Gaussian model of the spatial distance between image patches is adopted for the consideration of local and global contrast calculation. A new fusion method is designed to combine the feature maps for computing the final saliency map for stereoscopic images. Experimental results on a recent eye tracking database show the superior performance of the proposed method over other existing ones in saliency estimation for 3D images.
Yuming Fang 0001, Junle Wang, Manish Narwaria, Patrick Le Callet, Weisi Lin
VCIP4
2013 Visual Attention and Applications in Multimedia Technologies
abstract
Making technological advances in the field of human-machine interactions requires that the capabilities and limitations of the human perceptual system be taken into account. The focus of this report is an important mechanism of perception, visually selective attention, which is becoming more and more important for multimedia applications. We introduce the concept of visual attention and describe its underlying mechanisms. In particular, we introduce the concepts of overt and covert visual attention, and of bottom-up and top-down processing. Challenges related to modeling visual attention and their validation using ad hoc ground truth are also discussed. Examples of the usage of visual attention models in image and video processing are presented. We emphasize multimedia delivery, retargeting and quality assessment of image and video, medical imaging, and the field of stereoscopic 3-D image applications.
Patrick Le Callet, Ernst Niebur
Proc. IEEE1
2013 How Does Image Content Affect the Added Value of Visual Attention in Objective Image Quality Assessment?
abstract
Our previous research has demonstrated that adding natural scene saliency (NSS) obtained from eye-tracking data may improve an objective metric's performance in predicting perceived image quality. In this letter, we further investigate the image content dependency of this improvement. Results show that the variation in saliency between observers highly depends on image content, and that this variation predicts the extent to which a certain image may profit from adding saliency in the objective image quality assessment.
Hantao Liu, Ulrich Engelke, Junle Wang, Patrick Le Callet, Ingrid Heynderickx
IEEE Signal Process. Lett.4
2013 Comparative Study of Fixation Density Maps
abstract
Fixation density maps (FDM) created from eye tracking experiments are widely used in image processing applications. The FDM are assumed to be reliable ground truths of human visual attention and as such, one expects a high similarity between FDM created in different laboratories. So far, no studies have analyzed the degree of similarity between FDM from independent laboratories and the related impact on the applications. In this paper, we perform a thorough comparison of FDM from three independently conducted eye tracking experiments. We focus on the effect of presentation time and image content and evaluate the impact of the FDM differences on three applications: visual saliency modeling, image quality assessment, and image retargeting. It is shown that the FDM are very similar and that their impact on the applications is low. The individual experiment comparisons, however, are found to be significantly different, showing that inter-laboratory differences strongly depend on the experimental conditions of the laboratories. The FDM are publicly available to the research community.
Ulrich Engelke, Hantao Liu, Junle Wang, Patrick Le Callet, Ingrid Heynderickx, Hans-Jürgen Zepernick, Anthony J. Maeder
IEEE Trans. Image Process.4
2013 Computational Model of Stereoscopic 3D Visual Saliency
abstract
Many computational models of visual attention performing well in predicting salient areas of 2D images have been proposed in the literature. The emerging applications of stereoscopic 3D display bring an additional depth of information affecting the human viewing behavior, and require extensions of the efforts made in 2D visual modeling. In this paper, we propose a new computational model of visual attention for stereoscopic 3D still images. Apart from detecting salient areas based on 2D visual features, the proposed model takes depth as an additional visual dimension. The measure of depth saliency is derived from the eye movement data obtained from an eye-tracking experiment using synthetic stimuli. Two different ways of integrating depth information in the modeling of 3D visual attention are then proposed and examined. For the performance evaluation of 3D visual attention models, we have created an eye-tracking database, which contains stereoscopic images of natural content and is publicly available, along with this paper. The proposed model gives a good performance, compared to that of state-of-the-art 2D models on 2D images. The results also suggest that a better performance is obtained when depth information is taken into account through the creation of a depth saliency map, rather than when it is integrated by a weighting method.
Junle Wang, Matthieu Perreira Da Silva, Patrick Le Callet, Vincent Ricordel
IEEE Trans. Image Process.3
2013 Low-Cost Eye Gaze Prediction System for Interactive Networked Video Streaming
abstract
Eye gaze is now used as a content adaptation trigger in interactive media applications, such as customized advertisement in video, and bit allocation in streaming video based on region-of-interest (ROI). The reaction time of a gaze-based networked system, however, is lower-bounded by the network round trip time (RTT). Furthermore, only low-sampling-rate gaze data is available when commonly available webcam is employed for gaze tracking. To realize responsive adaptation of media content even under non-negligible RTT and using common low-cost webcams, we propose a Hidden Markov Model (HMM) based gaze-prediction system that utilizes the visual saliency of the content being viewed. Specifically, our HMM has two states corresponding to two of human's intrinsic gaze behavioral movements, and its model parameters are derived offline via analysis of each video's visual saliency maps. Due to the strong prior of likely gaze locations offered by saliency information, accurate runtime gaze prediction is possible even under large RTT and using common webcam. We demonstrate the applicability of our low-cost gaze prediction system by focusing on ROI-based bit allocation for networked video streaming. To reduce transmission rate of a video stream without degrading viewer's perceived visual quality, we allocate more bits to encode the viewer's current spatial ROI, while devoting fewer bits in other spatial regions. The challenge lies in overcoming the delay between the time a viewer's ROI is detected by gaze tracking, to the time the effected video is encoded, delivered and displayed at the viewer's terminal. To this end, we use our proposed low-cost gaze prediction system to predict future eye gaze locations, so that optimized bit allocation can be performed for future frames. Through extensive subjective testing, we show that bit-rate can be reduced by up to 29% without noticeable visual quality degradation when RTT is as high as 200 ms.
Yunlong Feng, Gene Cheung, Wai-tian Tan, Patrick Le Callet, Yusheng Ji
IEEE Trans. Multim.4
2012 Analysis and improvement of a paired comparison method in the application of 3DTV subjective experiment
abstract
Paired comparison is a frequently used method in psychophysical studies. However, with the increase of the number of the stimuli, the number of comparisons increases exponentially. Square design is one of the balanced sub-set paired comparison methods which could reduce the number of comparisons while producing comparably precise results under some assumptions. However, when there are observation errors from observers' attentiveness, the square design would produce large estimation errors. Thus, an improved square design which is robust to observation errors is proposed. Using a Monte Carlo simulation, the proposed method is evaluated and shows improvement in efficiency. The original design is applied in a visual discomfort subjective test of 3DTV. In addition, both of the two designs are studied by utilizing our previous full comparison data. The test results showed that the proposed improved square design is more robust to observation errors. Another important finding is that the influence of the occurrence of some other stimuli on voting is significant. Whether the proposed method could reduce the prediction errors induced by it is still under study.
Jing Li 0026, Marcus Barkowsky, Patrick Le Callet
ICIP3
2012 An edge-based structural distortion indicator for the quality assessment of 3D synthesized views
abstract
3D-TV applications require the generation of novel viewpoints through Depth-Image-Based-Rendering methods. These synthesized views need to be assessed by a reliable quality metric. Most of the proposed metrics are inspired from 2D commonly used quality metrics. Yet, the latter were originally designed to address 2D compression distortions which are different from the distortions related to DIBR processes. We propose an edge-based method that indicates the level of structural degradation in the synthesized image. The first results are encouraging since the correlation to subjective scores is higher than other tested metrics.
Emilie Bosc, Patrick Le Callet, Luce Morin, Muriel Pressigout
PCS2
2012 A Perceptually Relevant Channelized Joint Observer (PCJO) for the Detection-Localization of Parametric Signals
abstract
Many numerical observers have been proposed in the framework of task-based approach for medical image quality assessment. However, the existing numerical observers are still limited in diagnostic tasks: the detection task has been largely studied, while the localization task concerning one signal has been little studied and the localization of multiple signals has not been studied yet. In addition, most existing numerical observers need a priori knowledge about all the parameters of the underdetection signals, while only a few of them need at least two signal parameters. In this paper, we propose a novel numerical observer called the perceptually relevant channelized joint observer (PCJO), which cannot only detect but also localize multiple signals with unknown amplitude, orientation, size and location. We validated the PCJO for predicting human observer task performance by conducting a clinically relevant free-response subjective experiment in which six radiologists (including two experts) had to detect and localize multiple sclerosis (MS) lesions on magnetic resonance (MR) images. By using the jackknife alternative free-response operating characteristic (JAFROC) as the figure of merit (FOM), the detection-localization task performance of the PCJO was evaluated and then compared to that of the radiologists and two other numerical observers--channelized hotelling observer (CHO) and Goossenss CHO for detecting asymmetrical signals with random orientations. Overall, the results show that the PCJO performance was closer to that of the experts than to that of the other radiologists. The JAFROC1 FOMs of the PCJO (around 0.75) are not significantly different from those of the two experts (0.7672 and 0.7110), while the JAFROC1 FOMs of the numerical observers mentioned above (always over 0.84) outperform those of the experts. This indicates that the PCJO is a promising method for predicting radiologists' performance in the joint detection-localization task.
Lu Zhang 0037, Christine Cavaro-Ménard, Patrick Le Callet, Jean-Yves Tanguy
IEEE Trans. Medical Imaging3
2011 Can 3D synthesized views be reliably assessed through usual subjective and objective evaluation protocols?
abstract
This paper addresses the problem of evaluating virtual view synthesized images in the multi-view video context. As a matter of fact, view synthesis brings new types of distortion. The question refers to the ability of the traditional used objective metrics to assess synthesized views quality, considering the new types of artifacts. The experiments conducted to determine their reliability consist in assessing seven different view synthesis algorithms. Subjective and objective measurements have been performed. Results show that the most commonly used objective metrics can be far from human judgment depending on the artifact to deal with.
Emilie Bosc, Martin Köppel, Romuald Pépion, Muriel Pressigout, Luce Morin, Patrick Ndjiki-Nya, Patrick Le Callet
ICIP7
2011 A Subjective Evaluation of 3D Iptv Broadcasting Implementations Considering Coding and Transmission Degradation
abstract
This paper describes the results of a subjective test to assess current technology used for 3DTV broadcasting. As a first aspect, the performance of the currently deployed coding schemes was compared to state of the art algorithms. Our results show that down sampling and packing 3D stereoscopic videos according to the so called Side-By-Side format gives the highest perceived quality for a given bit rate. The second aspect of the study was to investigate how common 2D error concealment algorithms perform in case of 3D, and how their 3D-related performance compares with the 2D case. The results provide information on whether binocular suppression or binocular rivalries play the most important role for 3D video quality under transmission error. The results indicate that binocular rivalries and related visual discomfort are the dominant factors. Another aspect of the paper is a comparison of the test results with results from different labs to evaluate the repeatability of a subjective experiment in the 3D case, and to compare the employed test methodologies. Here, the study shows the variation between observers when they are rating visual discomfort and illustrates the difficulty to evaluate this new dimension.
Pierre R. Lebreton, Alexander Raake, Marcus Barkowsky, Patrick Le Callet
ISM4
2010 On the perceptual similarity of realistic looking tone mapped High Dynamic Range images
abstract
High Dynamic Range (HDR) images are usually displayed on conventional Low Dynamic Range (LDR) displays because of the limited availability of HDR displays. For the conversion of the large dynamic luminance range into the eight bit quantized values, parameterized Tone Mapping Operators (TMO) are applied. Human observers are able to optimize the parameters in order to get the highest Quality of Experience by judging the displayed LDR images on a realism scale. In the study presented in this paper, two TMOs with three parameters each were evaluated by observers in a subjective experiment. Although the chosen parameter settings vary largely, the chosen images appear to have the same QoE for the observers. In order to assess this similarity objectively, three commonly used image quality measurement algorithms were applied. Their agreement with the preference of the observers was analyzed and it was found that the Visual Difference Predictor (VDP) outperforms the Structural Similarity Index and the Root Mean Square Error. A threshold value for VDP is derived that indicates when two LDR images appear to have the same Quality of Experience.
Marcus Barkowsky, Patrick Le Callet
ICIP2
2010 Video quality assessment: From 2D to 3D - Challenges and future trends
abstract
Three-dimensional (3D) video is gaining a strong momentum both in the cinema and broadcasting industries as it is seen as a technology that will extensively enhance the user's visual experience. One of the major concerns for the wide adoption of such technology is the ability to provide sufficient visual quality, especially if 3D video is to be transmitted over a limited bandwidth for home viewing (i.e. 3DTV). Means to measure perceptual video quality in an accurate and practical way is therefore of highest importance for content providers, service providers, and display manufacturers. This paper discusses recent advances in video quality assessment and the challenges foreseen for 3D video. Both subjective and objective aspects are examined. An outline of ongoing efforts in standards-related bodies is also provided.
Quan Huynh-Thu, Patrick Le Callet, Marcus Barkowsky
ICIP2
2010 Linking distortion perception and visual saliency in H.264/AVC coded video containing packet loss
abstract
In this paper, distortions caused by packet loss during video transmission are evaluated with respect to their perceived annoyance. In this respect, the impact of visual saliency on the level of annoyance is of particular interest, as regions and objects in a video frame are typically not of equal importance to the viewer. For this purpose, gaze patterns from a task free eye tracking experiment were utilised to identify salient regions in a number of videos. Packet loss was then introduced into the bit stream such as that the corresponding distortions appear either in a salient region or in a non-salient region. A subjective experiment was then conducted in which human observers rated the annoyance of the distortions in the videos. The outcomes show a strong tendency that distortions in a salient region are indeed perceived as much more annoying as compared to distortions in the non-salient region. The saliency of the distorted image content was further found to have a larger impact on the perceived annoyance as compared to the distortion duration. The findings of this work are considered to be of great use to improve prediction performance of video quality metrics in the context of transmission errors.
Ulrich Engelke, Romuald Pépion, Patrick Le Callet, Hans-Jürgen Zepernick
VCIP3
2010 Overt visual attention for free-viewing and quality assessment tasks: Impact of the regions of interest on a video quality metric
Olivier Le Meur, Alexandre Ninassi, Patrick Le Callet, Dominique Barba
Signal Process. Image Commun.3
2010 Do video coding impairments disturb the visual attention deployment?
Olivier Le Meur, Alexandre Ninassi, Patrick Le Callet, Dominique Barba
Signal Process. Image Commun.3
2009 Region-of-Interest intra prediction for H.264/AVC error resilience
abstract
Packets in a video bitstream contain data with different levels of importance that yield unequal amounts of quality distortion when lost. In order to avoid sharp quality degradation due to packet loss, we propose in this paper an error resilience method that is applied to the region of interest (RoI) of the picture. This method protects the RoI while not yielding significant overhead. We perform an eye tracking test to determine the RoIs of a video sequence and we assess the performance of the proposed model in error-prone environments by means of a subjective quality test. Loss simulation results show that stopping the temporal error propagation in the RoIs of the pictures helps preserving an acceptable visual quality in the presence of packet loss.
Fadi Boulos, Wei Chen 0054, Benoît Parrein, Patrick Le Callet
ICIP4
2009 What we see is most likely to be what matters: Visual attention and applications
abstract
The computational modeling of the visual attention is receiving increasing attention from the computer vision community. Several bottom-up models have been proposed. In spite of their complexities, these models are still a basic description of our visual system. Review of resulting approaches of these efforts are presented in the first part of this paper. Limitations of these approaches are introduced and several research trends are given. Among them, the most important one might be the use of prior knowledge, conjointly with the low-level visual features. Concomitantly with visual attention (VA) modeling progress, the image and video processing community is increasingly considering VA models in different fields or services. Current and future applications of VA models are discussed in the second part.
Olivier Le Meur, Patrick Le Callet
ICIP2
2008 Using disparity for quality assessment of stereoscopic images
abstract
3DTV has been widely studied these last years from a technical point of view but the related perceived quality evaluations do not follow this hype. This paper firstly reviews quality assessment issues for 3DTV. Compared to 2D quality measure, depth information adds several new problems considering quality assessment. Nevertheless, efforts made for 2D content quality estimation can be used for an extension to 3D. In a second part, this paper proposes an adaptation of 2D metrics to 3D in the context of coding artifacts and stereoscopic images. Distortion on disparity is introduced to improve conventional 2D metrics. Performances have been evaluated using subjective tests.
Alexandre Benoît, Patrick Le Callet, Patrizio Campisi, Romain Cousseau
ICIP2
2008 Which semi-local visual masking model forwavelet based image quality metric?
abstract
Properties and models of the human visual system (HVS) are the fundaments for most of sufficient objective image or video quality metrics. Among HVS properties, visual masking is a sensitive issue. Many models exist in literature. Simplest models can only predict visibility threshold for very simple cue while for natural images one should consider more complex approaches such as semi-local masking. Our previous work has shown the positive impact of incorporating semi-local masking in image quality metric according to one subjective study. It is important to consolidate this work with different subjective experiments. In this paper, different visual masking models, including contrast masking and semi-local masking, are evaluated according to three subjective studies. These subjective experiments were conducted with different protocols, different types of display devices, different contents and different populations.
Alexandre Ninassi, Olivier Le Meur, Patrick Le Callet, Dominique Barba
ICIP3
2008 Objective quality assessment of color images based on a generic perceptual reduced reference
Mathieu Carnec, Patrick Le Callet, Dominique Barba
Signal Process. Image Commun.2
2007 Does where you Gaze on an Image Affect your Perception of Quality? Applying Visual Attention to Image Quality Metric
abstract
The aim of an objective image quality assessment is to find an automatic algorithm that evaluates the quality of pictures or video as a human observer would do. To reach this goal, researchers try to simulate the Human Visual System (HVS). Visual attention is a main feature of the HVS, but few studies have been done on using it in image quality assessment. In this work, we investigate the use of the visual attention information in their final pooling step. The rationale of this choice is that an artefact is likely more annoying in a salient region than in other areas. To shed light on this point, a quality assessment campaign has been conducted during which eye movements have been recorded. The results show that applying the visual attention to image quality assessment is not trivial, even with the ground truth.
Alexandre Ninassi, Olivier Le Meur, Patrick Le Callet, Dominique Barba
ICIP (2)3
2007 Impact of the Resolution on the Difference of Perceptual Video Quality Between CRT and LCD
abstract
The incoming of the high-definition new visual experience at home has boosted the new display technologies such as liquid crystal displays (LCD), plasma and projectors. These technologies enable the increase of the screen size necessary to sense a cinema-like experience. However, they introduce some new visual shortcomings not present with the mature CRT technology. In this paper, some subjective tests are described which highlight a difference of the perceptual video quality between CRT and LCD. Moreover, it's observed that this loss of quality on LCD is more important with high resolution sequences than with standard resolution ones. This influence of the resolution is particularly explainable in the case of the LCD motion blur defect.
Sylvain Tourancheau, Patrick Le Callet, Dominique Barba
ICIP (3)2
2007 A robust image watermarking technique based on quantization noise visibility thresholds
Florent Autrusseau, Patrick Le Callet
Signal Process.2
2006 Efficient Saliency-Based Repurposing Method
abstract
Images play a very relevant role in our daily life. People now can easily shoot and share pictures thanks to the exponential growth of the portable medias, such as digital cameras, mobile phone. As the display size of those devices is relatively small, browsing large pictures remains difficult. Content re-purposing is an elegant solution to deal with this problem. It consists in cropping the images in order to display only the most interesting parts of the picture. A new algorithm is proposed in this paper; the experiments described herein, leading to a qualitative and a quantitative assessment, show that the proposed solution outperforms the conventional method.
Olivier Le Meur, Xavier Castellani, Patrick Le Callet, Dominique Barba
ICIP3
2006 From SD to HD Television: Effects of H.264 Distortions Versus Display Size on Quality of Experience
abstract
High Definition Television (HDTV) is the new broadcasting system designed to take the place of Standard Definition Television (SDTV) at home in the near future. This system requires modification of many features in the broadcasting chain with an overall objective of reaching a noticeably higher quality of experience. Since broadcasters desire a high level of service acceptability, they require efficient measurements of quality of experience. The purpose of this paper is to provide such measurements concerning the noticeable artifacts in H.264 distortions over a range of display sizes and comparing HDTV to SDTV. A subjective characterization of some HDTV quality of experience aspects is proposed and the results are discussed.
Stéphane Péchard, Mathieu Carnec, Patrick Le Callet, Dominique Barba
ICIP3
2006 A Coherent Computational Approach to Model Bottom-Up Visual Attention
abstract
Visual attention is a mechanism which filters out redundant visual information and detects the most relevant parts of our visual field. Automatic determination of the most visually relevant areas would be useful in many applications such as image and video coding, watermarking, video browsing, and quality assessment. Many research groups are currently investigating computational modeling of the visual attention system. The first published computational models have been based on some basic and well-understood Human Visual System (HVS) properties. These models feature a single perceptual layer that simulates only one aspect of the visual system. More recent models integrate complex features of the HVS and simulate hierarchical perceptual representation of the visual input. The bottom-up mechanism is the most occurring feature found in modern models. This mechanism refers to involuntary attention (i.e., salient spatial visual features that effortlessly or involuntary attract our attention). This paper presents a coherent computational approach to the modeling of the bottom-up visual attention. This model is mainly based on the current understanding of the HVS behavior. Contrast sensitivity functions, perceptual decomposition, visual masking, and center-surround interactions are some of the features implemented in this model. The performances of this algorithm are assessed by using natural images and experimental measurements from an eye-tracking system. Two adequate well-known metrics (correlation coefficient and Kullbacl-Leibler divergence) are used to validate this model. A further metric is also defined. The results from this model are finally compared to those from a reference bottom-up model.
Olivier Le Meur, Patrick Le Callet, Dominique Barba, Dominique Thoreau
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 A Convolutional Neural Network Approach for Objective Video Quality Assessment
abstract
This paper describes an application of neural networks in the field of objective measurement method designed to automatically assess the perceived quality of digital videos. This challenging issue aims to emulate human judgment and to replace very complex and time consuming subjective quality assessment. Several metrics have been proposed in literature to tackle this issue. They are based on a general framework that combines different stages, each of them addressing complex problems. The ambition of this paper is not to present a global perfect quality metric but rather to focus on an original way to use neural networks in such a framework in the context of reduced reference (RR) quality metric. Especially, we point out the interest of such a tool for combining features and pooling them in order to compute quality scores. The proposed approach solves some problems inherent to objective metrics that should predict subjective quality score obtained using the single stimulus continuous quality evaluation (SSCQE) method. This latter has been adopted by video quality expert group (VQEG) in its recently finalized reduced referenced and no reference (RRNR-TV) test plan. The originality of such approach compared to previous attempts to use neural networks for quality assessment, relies on the use of a convolutional neural network (CNN) that allows a continuous time scoring of the video. Objective features are extracted on a frame-by-frame basis on both the reference and the distorted sequences; they are derived from a perceptual-based representation and integrated along the temporal axis using a time-delay neural network (TDNN). Experiments conducted on different MPEG-2 videos, with bit rates ranging 2-6 Mb/s, show the effectiveness of the proposed approach to get a plausible model of temporal pooling from the human vision system (HVS) point of view. More specifically, a linear correlation criteria, between objective and subjective scoring, up to 0.92 has been obtained on a set of typical TV videos.
Patrick Le Callet, Christian Viard-Gaudin, Dominique Barba
IEEE Trans. Neural Networks1
2005 Flexible Storage of Still Images with a Perceptual Quality Criterion
Vincent Ricordel, Patrick Le Callet, Mathieu Carnec, Benoît Parrein
ACIVS2
2005 Visual features for image quality assessment with reduced reference
abstract
This paper focuses on the definition and selection of features to design a reduced description image quality assessment method. For such methods, the main problem is the choice of the features which constitute this reduced description. The best features are the ones which enable to produce the highest correlation between produced quality scores and subjective ones. So we test several feature types and measure their impact for image quality assessment. We show that the use of structural information and the combination of different feature types permit to get high performances.
Mathieu Carnec, Patrick Le Callet, Dominique Barba
ICIP (1)2
2005 A spatio-temporal model of the selective human visual attention
abstract
A new spatio-temporal model for simulating the bottom-up visual attention is proposed. It has been built from numerous important properties of the human visual system (HVS). This paper focuses both on the architecture of the model and on its performances. Given that the spatial model of the bottom-up visual attention has already been defined [O. Le Meur et al., 2004], the temporal dimension is more accurately described. A qualitative and quantitative comparison with human fixations collected from an eye tracking apparatus is undertaken. From the former, the quality of the prediction is deemed very good whereas the latter illustrates that the best predictor of the human fixation consists of the sum all visual features (achromatic, chromatic and motion).
Olivier Le Meur, Dominique Thoreau, Patrick Le Callet, Dominique Barba
ICIP (3)3
2004 Performance assessment of a visual attention system entirely based on a human vision modeling
abstract
It is now commonly assumed that the human visual attention, which is a selecting process of the most relevant locations in a scene according to a particular behavior, is driven by both top-down (task-dependent) and bottom-up (signal-dependent) control. A new model attempting to simulate the bottom-up process has been designed Le Meur, O et al., (2004). This model is purely based on visual system properties that provides noticeable advantages compared to the classical published approaches. This paper focuses on the performance assessment of this model by achieving a comparison with real fixation points stemming from eye-tracking apparatus both subjectively and objectively.
Olivier Le Meur, Patrick Le Callet, Dominique Barba, Dominique Thoreau
ICIP2
2003 A robust quality metric for color image quality assessment
abstract
In this paper, we propose a visual color image quality metric with full reference image for the evaluation of coding schemes. This metric is based on human visual system properties in order to get the best correspondence with human judgements. Contrary to some others objective criteria, it doesn't use any information on the type of degradations introduced by coding schemes. We use two main stages: the first one in order to compute visual representation of images (based on results from psychophysics experiments conducted in our laboratory) and the second in order to pool errors between visual representation of two images. We propose a new approach for this pooling stage based on the density of errors and their structure. We also show the interest of such pooling method. We compare results of the metric with human judgments on a database of 140 images distorted with 3 types of compression schemes (JPEG, JPEG2000 and a ROI-based algorithm). High performances are obtained leading to assure that the metric is robust, so this approach constitutes an alternative useful tool to PSNR for coding image searchers.
Patrick Le Callet, Dominique Barba
ICIP (1)1
2003 An image quality assessment method based on perception of structural information
abstract
This paper presents a new method to evaluate the quality if distorted images. This method is based on a comparison between the structural information extracted from the distorted image and from the original image. The interest of our method is that it uses reduced references containing perceptual structural information. First, a quick overview of image quality evaluation methods is given. Then the implementation of our human visual system (HVS) model is detailed. At last, results are given for quality evaluation of JPEG and JPEG2000 coded images. They show that our method provides results which are highly correlated with human judgments (mean opinion score). This method has been implemented in an application available on the Internet.
Mathieu Carnec, Patrick Le Callet, Dominique Barba
ICIP (3)2
2003 Robust approach for color image quality assessment
Patrick Le Callet, Dominique Barba
VCIP1
2003 New perceptual quality assessment method with reduced reference for compressed images
Mathieu Carnec, Patrick Le Callet, Dominique Barba
VCIP2