EDBT 2026 Demo / reviewers in the wild / expert
Mohamed-Chaker Larabi
dblp:31/5962 · also Chaker Larabi
· DBLP profile ↗
84ranked-venue papers
1as first author
27since 2021 · last 2026
0000-0003-4511-5381ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 78 · 1 first-author · 24 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Audio-driven visual attention for lightweight audio-visual quality assessmentabstractMultimedia has become an integral component of the Internet and modern digital experiences, playing a vital role in our everyday lives. The majority of content that we consume daily combines both audio and video modalities. This emphasizes the demand for an effective audio-visual quality assessment. However, current state-of-the-art audio-visual quality assessment methods often require a high computational cost and execution time. Moreover, they ignore important factors, such as the signal’s context, that influence what users expect in terms of quality across different circumstances. In this study, we therefore introduce an approach that addresses these problems. Inspired by the complementary nature of audio and vision, our methods leverage audio cues to selectively amplify visually informative regions, allowing the model to capture the cross-modal contextual dependencies in a computationally efficient manner. Experimental results show that our proposed model significantly reduces the computational complexity by more than 80% of both parameter size and runtime while still providing a favorable performance compared to state-of-the-art methods. The source code of this work will be made publicly available. Ha Thu Nguyen, Seyed Ali Amirshahi, Katrien De Moor, Mohamed-Chaker Larabi |
Neurocomputing | 4 |
| 2026 | TransformAR: A light-weight transformer-based metric for Augmented Reality quality assessmentabstractAs Augmented Reality (AR) technology continues to gain traction in various sectors, ensuring a superior user experience has become an essential challenge for both academic researchers and industry professionals. However, the task of automatically predicting the quality of AR images remains difficult due to several inherent challenges, particularly the issue of visual confusion arising from the overlap of virtual and real-world elements. This paper introduces transformAR, a novel and efficient transformer-based framework designed to objectively assess the quality of AR images. The proposed model uses pre-trained vision transformers to capture content features from AR images, calculates distance vectors to measure the impact of distortions, and employs cross-attention-based decoders to effectively model the perceptual qualities of the AR images. Additionally, the training framework uses regularization techniques and label smoothing-like method to reduce the risk of overfitting. Through comprehensive experiments, we demonstrate that transformAR outperforms existing state-of-the-art approaches, offering a more reliable and scalable solution for AR image quality assessment. Aymen Sekhri, Mohamed-Chaker Larabi, Seyed Ali Amirshahi |
Signal Process. Image Commun. | 2 |
| 2026 | Enhancing Content Representation for AR Image Quality Assessment Using Knowledge DistillationabstractAugmented Reality (AR) is a major immersive media technology that enriches our perception of reality by overlaying digital content (the foreground) onto physical environments (the background). It has far-reaching applications, from entertainment and gaming to education, healthcare, and industrial training. Nevertheless, challenges such as visual confusion and classical distortions can result in user discomfort when using the technology. Evaluating AR quality of experience becomes essential to measure user satisfaction and engagement, facilitating the refinement necessary for creating immersive and robust experiences. Though the scarcity of data and the distinctive characteristics of AR technology render the development of effective quality assessment metrics challenging. This paper presents a deep learning-based objective metric designed specifically for assessing image quality for AR scenarios. The approach entails four key steps, (1) fine-tuning a self-supervised pre-trained vision transformer to extract prominent features from reference images and distilling this knowledge to improve representations of distorted images, (2) quantifying distortions by computing shift representations, (3) employing cross-attention-based decoders to capture perceptual quality features, and (4) integrating regularization techniques and label smoothing to address the overfitting problem. To validate the proposed approach, we conduct extensive experiments on the ARIQA dataset. The results showcase the superior performance of our proposed approach across all model variants, namely TransformAR, TransformAR-KD, and TransformAR-KD+ in comparison to existing state-of-the-art methods. Aymen Sekhri, Seyed Ali Amirshahi, Mohamed-Chaker Larabi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | ICE-Cubed: Inpainting of Cinematographic Elements for Intelligent Context Expansion
David Traparic, Maugan De Murcia, Mohamed-Chaker Larabi, Ladjel Bellatreche |
ACIVS | 3 |
| 2025 | Lightweight Image Quality Prediction Guided by Perceptual Ranking FeedbackabstractAutomatic Image Quality Assessment (IQA) remains a difficult challenge due to the complexity of mimicking the Human Visual System (HVS) and the limitations of traditional objective Image Quality Metrics (IQM). Existing learnable methods often involve high computational costs and fail to adequately capture the nuanced perceptual characteristics of the HVS, including the human ability to rank image quality and human sensitivity to differences in areas with high-frequency. In this study, we propose an effective approach that addresses these challenges by incorporating the characteristics of HVS and the perceptual classification into a lightweight IQM framework based on the transformer architecture. This allows our method to capture long-range dependencies effectively. Our approach leverages Objective Error Maps (OEMs) to enhance sensitivity to visual errors and employs a ranking module as an objective function, providing feedback on the perceptual quality at the feature level. Experimental results demonstrate that our approach not only achieves competitive performance compared to state-of-the-art IQMs but also significantly reduces computational complexity. Aymen Sekhri, Mohamed-Chaker Larabi, Seyed Ali Amirshahi |
ICASSP | 2 |
| 2025 | ARaBIQA: A Novel Blind Image Quality Assessment Model for Augmented RealityabstractEnsuring the quality of Augmented Reality (AR) experiences is crucial for achieving user satisfaction in many applications such as navigation, education, and healthcare. However, automatic AR quality assessment is challenging due to limited data and the lack of a reference image notion in real-world scenarios. Hence, blind quality assessment appears to be the only plausible solution. Existing blind IQA metrics often struggle to capture perceptual features in AR content as effectively as they do in natural images. We propose ARaBIQA, the first blind image quality assessment (BIQA) method designed specifically for AR content. Using a self-supervised approach, ARaBIQA learns low-level AR-specific features, including distortions and visual confusion, and combines them with high-level content features through a joint fine-tuning strategy to produce robust quality predictions. The experimental results show that ARaBIQA outperforms existing blind IQA metrics, and ablation studies further validate its effectiveness. Aymen Sekhri, Mohamed-Chaker Larabi, Seyed Ali Amirshahi |
ICIP | 2 |
| 2025 | Latent Space Stability vs. Perceptual Sensitivity: A Study of Visual Encoders under DistortionabstractRobust and distortion-aware visual representations are paramount for perceptual image quality assessment (IQA) and downstream visual understanding under real-world degradations. In this paper, we conduct a comprehensive analysis of state-of-the-art visual encoders, including CLIP, DINO, ConvNeXt, EfficientNet, and ResNet, under common distortions such as Gaussian blur, motion blur, compression artifacts, and Gaussian noise. We employ latent feature divergence, ANOVA-based effect size analysis, and dimension-wise mean absolute differences (MAD) to assess model robustness and sensitivity. Our results reveal how architectural choices, training objectives, and data diversity shape a model’s ability to encode distortions within its latent space. These findings bridge representation learning and perceptual quality modeling, offering new insights into the development of distortion-resilient encoders for IQA. Abderrezzaq Sendjasni, Mohamed-Chaker Larabi |
MMSP | 2 |
| 2025 | Local Structure Matters: A Graph-Based Approach to Point Cloud Perceptual Quality AssessmentabstractThis paper introduces a novel graph-based framework for point cloud quality assessment (PCQA) that bridges the gap between objective metrics and perceptual quality. The framework constructs local graphs around key points identified through curvature analysis and adaptively incorporates spatial and color information to model local structures effectively. The proposed approach captures structural and visual characteristics by leveraging signal processing on graphs and extracting domain-specific features from the geometry, appearance, and spectral domains. Extensive evaluations using SJTU-PCQA and WPC, two publicly available datasets, demonstrate the superiority of the proposed framework, achieving state-of-the-art performance and setting new benchmarks among graph-based approaches for PCQA. Furthermore, the ablation study showcases the complementary benefits of integrating domain-specific features, highlighting the robustness and adaptability of the proposed method across various point cloud densities and qualities. Abderrezzaq Sendjasni, Mohamed-Chaker Larabi |
QoMEX | 2 |
| 2025 | HDR Video Composition using Differently Exposed Stereo LDR VideosabstractHigh Dynamic Range (HDR) videos provide a wider luminance range and more accurate real-world illumination compared to Low Dynamic Range (LDR) videos. Unlike images, HDR video generation requires additional attention to ensure flicker-free results while preserving temporal coherence. This paper introduces a deep learning-based approach to generate HDR videos from stereo LDR frames with different exposures, effectively reducing flickering and artifacts. It employs a two-stage pipeline with individual convolutional neural networks. At each time step, the HDR video frame is constructed by combining its corresponding stereo LDR frames with the previous HDR frame. This strategy enhances temporal coherence by ensuring smoother transitions between consecutive frames. The process begins by recovering high dynamic range radiance maps from the initial stereo LDR frames, which are subsequently replaced by the composed HDR frames in later iterations. Experiments on existing datasets demonstrate the effectiveness of our approach in generating high-quality HDR videos with improved temporal consistency. Shashaank Aswatha Mattur, Mohamed-Chaker Larabi |
VCIP | 2 |
| 2025 | Embedding similarity guided license plate super resolutionabstractSuper-resolution (SR) techniques play a pivotal role in enhancing the quality of low-resolution images, particularly for applications such as security and surveillance, where accurate license plate recognition is crucial. This study proposes a novel framework that combines pixel-based loss with embedding similarity learning to address the unique challenges of license plate super-resolution (LPSR). The introduced pixel and embedding consistency loss (PECL) integrates a Siamese network and applies contrastive loss to force embedding similarities to improve perceptual and structural fidelity. By effectively balancing pixel-wise accuracy with embedding-level consistency, the framework achieves superior alignment of fine-grained features between high-resolution (HR) and super-resolved (SR) license plates. Extensive experiments on the CCPD and PKU dataset validate the efficacy of the proposed framework, demonstrating consistent improvements over state-of-the-art methods in terms of PSNR, SSIM, LPIPS, and optical character recognition (OCR) accuracy. These results highlight the potential of embedding similarity learning to advance both perceptual quality and task-specific performance in extreme super-resolution scenarios. Abderrezzaq Sendjasni, Mohamed-Chaker Larabi |
Neurocomputing | 2 |
| 2024 | Enhancing Perceptual Quality Assessment for 360-Degree Images Based on Adaptive Patch Labeling and Multi-Label LearningabstractThis paper delves into the intricate field of perceptual quality assessment specifically tailored for 360-degree images, aiming to advance the precision of quality models. In contrast to conventional methodologies that associate different regions within images to mean opinion scores (MOS), our study introduces a paradigm shift. We propose a novel approach where the model is trained to predict multi-labels derived from subjective and objective measures, leveraging a sophisticated quality labeling framework designed to capture nuanced perceptual distinctions across diverse regions in panoramic content. This allows for a flexible and stable training process. In addition, a loss function taking into account the magnitude and direction of quality labels under a multi-label learning scheme is designed, using absolute and directional distances as loss functions. Experimental results underscore the limitations of relying solely on MOS as unique labels. The efficiency of our approach becomes evident through improved performance, showcasing its potential for advancing the precision of perceptual quality assessment models. Abderrezzaq Sendjasni, Mohamed-Chaker Larabi |
ICIP | 2 |
| 2024 | Towards Light-Weight Transformer-Based Quality Assessment Metric for Augmented RealityabstractWith the rise of Augmented Reality (AR) technology, which enhances the real world by overlaying computer-generated content, immersive experiences are being offered in education, entertainment, healthcare, … Assessing the quality of AR scenarios is crucial for understanding and improving user satisfaction and engagement. However, developing objective AR quality assessment methods is challenging due to the lack of data and the inherent complexity of technology, particularly in the presence of visual confusion. Existing convolution neural network-based approaches suffer from limited receptive fields and are not effective at capturing global information in visually confused AR scenarios. Additionally, to the best of our knowledge, exploring transformer capabilities for AR quality assessment is missing. Therefore, this study introduces transformAR, a lightweight transformer-based model for objective quality assessment in AR applications. This approach leverages pretrained vision transformer-based encoders to capture image content information, computes distance vectors to quantify distortions, and employs cross-attention-based decoders to model perceptual quality features. The model also integrates adapted regularization techniques and label smoothing to mitigate overfitting. Experimental results demonstrate the effectiveness of transformAR, outperforming the few existing state-of-the-art methods. Aymen Sekhri, Seyed Ali Amirshahi, Mohamed-Chaker Larabi |
MMSP | 3 |
| 2024 | Embedding Similarity Learning for Extreme License Plate Super-ResolutionabstractSuper-resolution (SR) techniques play a crucial role in enhancing the quality of low-resolution images, with significant applications in fields such as security and surveillance, where license plate recognition is critical. This paper focuses on optimizing the super-resolution of license plates using embedding similarity learning. We proposed a novel framework that integrates a Siamese network with a super-resolution model to guide the SR model into enhancing the perceptual quality of reconstructed license plates. By leveraging embedding similarity through Contrastive loss, our approach ensures that the super-resolved images are perceptually and structurally closer to the original ones. The experiments on a synthetic dataset demonstrated that the proposed method outperforms traditional techniques that rely solely on pixel-based loss functions such as MSE. The introduction of embedding similarity loss significantly improves the PSNR and LPIPS metrics, in addition to the optical characters recognition rate. Abderrezzaq Sendjasni, Mohamed-Chaker Larabi |
MMSP | 2 |
| 2024 | CNN-LPQ: convolutional neural network combined to local phase quantization based approach for face anti-spoofing
Mebrouka Madi, Mohammed Khammari, Mohamed-Chaker Larabi |
Multim. Tools Appl. | 3 |
| 2023 | Self Patch Labeling Using Quality Distribution Estimation for CNN-Based 360-IQA TrainingabstractIn this study, we propose a methodology for estimating quality score distribution (QSD) for 360-IQA patch labeling. A collection of 2D-IQA models is used to generate a QSD for patches, inspired by how subjective quality ratings are gathered and handled. The proposed framework is first benchmarked on a subjectively annotated dataset, namely KonPatch-32k, in terms of patch quality classification. The best composition of QSD is then used to derive quality labels for patches sampled from 360-degree images. Furthermore, the quality labels are used in a multi-regression training strategy of CNN models. The ResNet-50 and EfficientNet-B5 are used to test the effectiveness of the proposed labeling framework on two publicly available 360-IQA datasets, namely OIQA and MVAQD. The experimental results demonstrated the efficacy of jointly using local and global qualities. The multi-regression proved to be a bit challenging on OIQA compared to MVAQD, reflecting the necessity to accurately regulate the training process. Abderrezzaq Sendjasni, Mohamed-Chaker Larabi |
ICIP | 2 |
| 2023 | Construction of a Video Inpainting Dataset Based on a Subjective StudyabstractVideo inpainting, the automated process of reconstructing missing or corrupted regions in video sequences, has gained significant attention in the fields of computer vision and image processing in recent years. However, a recurring remark in the literature has been the lack of a dedicated database specifically designed for video inpainting. As a result, existing inpainting studies have relied on locally created videos or datasets primarily intended for other applications. To address this limitation, this paper introduces the first publicly available video inpainting dataset accompanied by subjective scores, which closely aligns with real-world applications. The dataset covers three key inpainting scenarios: video hole completion, object removal, and post-stabilization inpainting. By providing this dataset, our goal is to facilitate the comparison and evaluation of both current and future video inpainting techniques. Moreover, we anticipate that it will serve as a solid foundation for the development of novel video inpainting assessment metrics in the future, thereby encouraging further advancements in this field. It is worth noting that, to the best of our knowledge, there is only one metric dedicated to video inpainting quality assessment, apart from those developed for image inpainting that can potentially be extended to videos. The dataset and the subjective scores are available on this link: https://github.com/rezkimed/VID-SS. Amine Mohamed Rezki, Mohamed-Chaker Larabi, Amina Serir |
MMSP | 2 |
| 2023 | Adaptive Patch Labeling and Multi-Label Feature Selection for 360-Degree Image Quality AssessmentabstractAssessing the quality of 360-degree images based on individual regions presents a challenging task. The lack of ground truth opinion scores (MOS) for specific regions makes it difficult to evaluate image quality accurately. Existing datasets only provide MOS for entire 360-degree images, which limits the granularity of assessment. To overcome this challenge, we propose a novel framework that employs adaptive patch labeling techniques. We leverage a set of 2D-IQA methods to generate quality score distributions for each patch in the 360-degree images. These distributions, combined with the available MOS, serve as labels for individual patches, providing a more comprehensive characterization of patch quality. Furthermore, we use these labels to adaptively select and refine deep neural features. By selectively choosing label-specific features, we enhance the accuracy and effectiveness of patch-based 360-degree image quality assessment. This approach allows us to focus on the most relevant and informative features for each patch, resulting in improved assessment performance. The experimental results on two benchmark datasets demonstrate that adaptive patch labeling and feature selection achieve accurate and reliable performances, thus advancing the field of 360-degree image quality assessment. Abderrezzaq Sendjasni, Mohamed-Chaker Larabi, Seif-Eddine Benkabou |
MMSP | 2 |
| 2023 | Towards Automatic Content Generation for Immersive Cinema Theater Based on Artificial IntelligenceabstractImmersive display systems like the one proposed by ICE® technology aims to enhance visual immersion by widening the field of view. However, creating immersive content while maintaining immersion integrity is a challenging task due to the sensitivity of human peripheral vision to flickering and movement. Moreover, identifying elements in videos that may disrupt immersion and determining whether they can be expanded into an immersive context is a complex and time-consuming process due to the lack of automatic methodologies. In this paper, we propose a pipeline for automatically generating content for lateral displays from movies. The pipeline consists of several steps. Firstly, the input content is divided into cinematic shots, and then further segmented into snippets. Next, domain-specific features are extracted using dedicated video deep learning models. Additionally, handcrafted features are computed to provide task-specific information. These extracted features are utilized to predict the required processing steps for generating lateral content that aligns with ground-truth annotations provided by cinema experts. The results obtained from our pipeline show promising accuracy and demonstrate the potential for this specialized application. David Traparic, Mohamed-Chaker Larabi, Ladjel Bellatreche |
MMSP | 2 |
| 2023 | Towards explainable deep visual saliency models
Sai Phani Kumar Malladi, Jayanta Mukhopadhyay, Mohamed-Chaker Larabi, Santanu Chaudhury |
Comput. Vis. Image Underst. | 3 |
| 2022 | Lighter and Faster Two-Pathway CMRNet for Video Saliency PredictionabstractExisting dynamic saliency prediction models face challenges like inefficient spatio-temporal feature integration, ineffective multi-scale feature extraction, and lacking domain adaptation because of huge pre-trained backbone networks. In this paper, we propose a two pathway architecture with effective feature integration of spatial and temporal domains at multiple scales for video saliency prediction. Frame and optical flow pathways extract features from video frame and optical flow maps, respectively using a series of cross-concatenated multi-scale residual (CMR) blocks. We name this network as two-pathway CMRNet (TP-CMRNet). Every CMR block follows a feature fusion and attention module for merging features from two pathways and guiding the network to weigh salient regions, respectively. A bi-directional LSTM module is used for learning the task by looking at previous and next video frames. We build a simple decoder for feature reconstruction into the final attention map. TP-CMRNet is comprehensively evaluated using three benchmark datasets: DHF1K, Hollywood-2, and UCF sports. We observe that our model performs at par with other deep dynamic models. In particular, we outperform all the other models with a lesser number of model parameters and lower inference time. Sai Phani Kumar Malladi, Jayanta Mukhopadhyay, Mohamed-Chaker Larabi, Santanu Chaudhury |
ICIP | 3 |
| 2022 | Transfer Learning from Vision Transformers or ConvNets for 360-Degree Images Quality AssessmentƒabstractCurrently, there are debates on the accuracy of vision transformers (ViTs) compared to ConvNets for image processing tasks. Image quality assessment (IQA) and particularly 360-IQA is lacking insights regarding their performances and robustness compared to the widely used ConvNets. This paper aims to investigate transfer learning from two pre-trained versions of ViTs and two ConveNets (ResNet-50 and EfficientNet-B3) for 360-degree image quality assessment with a focus on (i) the prediction accuracy and generalization ability and (ii) their adaptation to the specific characteristics of 360-degree images. Furthermore, the influence of adaptive patches sampling compared to simply using equirectangular content is analyzed with each architecture. Experimental findings on publicly available datasets (OIQA, CVIQ and MVAQD) show the superiority of ResNet-50 over ViTs and EfficientNet-B3 while requiring less computational time. Also, the base version of ViTs outperforms the larger one. Finally, except for CVIQ, both ViTs and ConveNets benefit from the adaptive sampling strategy, depicting the interest of taking 360-degree characteristics into account. Abderrezzaq Sendjasni, Mohamed-Chaker Larabi |
ICIP | 2 |
| 2022 | Investigating Normalization Methods for CNN-Based Image Quality AssessmentabstractPrior to training convolutional neural networks (CNNs) for image quality assessment (IQA), input normalization is sometimes recommended and sometimes not, according to the literature. Although input normalization is known to improve model training and helps in learning important features, it may result in the loss of information such as contrast, color, and luminance. To better explore this issue, we conduct an empirical study to first investigate the effect of normalization on model performance and then which normalization method best fits IQA among existing methods. The performances of the selected methods are statistically compared with three basic scaling methods. The application of normalization is found to be statistically significant on three IQA databases. The performance improvement on the overall databases, as well as per-individual degradation, is demonstrated in the experimental results. Abderrezzaq Sendjasni, David Traparic, Mohamed-Chaker Larabi |
ICIP | 3 |
| 2022 | Convolutional Neural Networks for Omnidirectional Image Quality Assessment: A BenchmarkabstractIn this paper, we conduct an extensive study on the use of pre-trained convolutional neural networks (CNNs) for omnidirectional image quality assessment (IQA). To cope with the lack of available IQA databases, transfer learning from seven pre-trained CNN models is investigated over retraining on standard 2D databases. In addition, we explore the influence of various image representations and training strategies on the model’s performance. A comparison of the use of projected versus radial content, and multichannel CNN versus patch-wise training is also covered. The experimental results on two publicly available databases are used to draw conclusions about which strategy best fits the visual quality prediction and at which computational cost. The analysis shows that retraining CNN models on 2D IQA databases improves the prediction accuracy. The latter and the required computational time are found to be significantly affected by the training strategy. Cross-database evaluations demonstrate that the nature and variety of the content impact the generalization ability of the models. Finally, we show that conclusions coming from other image processing communities may not hold for IQA. The provided discussion shall provide insights and recommendations when using pre-trained CNNs for omnidirectional IQA. Abderrezzaq Sendjasni, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Deep High Dynamic Range Imaging Using Differently Exposed Stereo ImagesabstractHigh dynamic range (HDR) image formation from low dynamic range (LDR) images of different exposures is a well researched topic in the past two decades. However, most of the developed techniques consider differently exposed LDR images that are acquired from the same camera view point, which assumes the scene to be static long enough to capture multiple images. In this paper, we propose to address the problem of HDR imaging from differently exposed LDR stereo images using an encoder-decoder based convolutional neural network (CNN). The proposed technique does not require the LDR stereo images to be explicitly rectified and disparity corrected before merging to HDR image, unlike conventional stereo matching methods. For training and evaluation, we consider an existing benchmark dataset of HDR stereo images. The experiments have shown some interesting results in comparison to the state-of-the-art approaches. The proposed end-to-end network is found to perform equally well on LDR images that are obtained from both stereo framework and single viewpoint. Shashaank M. Aswatha, Mohamed-Chaker Larabi |
ICIP | 2 |
| 2021 | Lighter and Faster Cross-Concatenated Multi-Scale Residual Block Based Network for Visual Saliency PredictionabstractExisting deep architectures for visual saliency prediction face problems like inefficient feature encoding, larger inference times, and a huge number of model parameters. One possible solution is to make the local and global contextual feature extraction computationally less intensive by a novel lighter architecture. In this work, we propose an end-to-end learnable, inter-scale information sharing residual block based architecture for saliency prediction. A series of these blocks are used for efficient multi-scale feature extraction followed by a dilated inception module (DIM) and a novel decoder. We name this network as cross-concatenated multi-scale residual (CMR) block based network, CMRNet. We comprehensively evaluate our architecture on three datasets: SALICON, MIT1003, and MIT300. Experimental results show that our model works at par with other state-of-the-art models. Especially, our model outperforms all the other models with a smaller inference time and a lesser number of model parameters. Sai Phani Kumar Malladi, Jayanta Mukhopadhyay, Mohamed-Chaker Larabi, Santanu Chaudhury |
ICIP | 3 |
| 2021 | Perceptually-Weighted Cnn For 360-Degree Image Quality Assessment Using Visual Scan-Path And JndabstractImage quality assessment of immersive content and more specifically 360-degree one is still in its infancy. There are many challenges regarding sphere vs. projected representation, human visual system (HVS) properties in a 360-degree environment, etc. In this paper, we propose the use of CNNs to design a no reference model to predict visual quality of 360-degree images. Instead of feeding the CNN with ERPs, visually important viewports are extracted based on visual scan-path prediction and given to a multi-channel CNN using DenseNet-121. Moreover, information about visual fixations and just noticeable difference are used to account for the HVS properties and make the network closer to human judgment. The scan-path is also used to create multiple instances of the database so as to perform a robust generalization analysis and compensate for the lack of databases. Abderrezzaq Sendjasni, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh |
ICIP | 2 |
| 2021 | Convolutional Neural Networks for Omnidirectional Image Quality Assessment: Pre-Trained or Re-Trained?abstractThe use of convolutional neural networks (CNN) for image quality assessment (IQA) becomes many researcher’s focus. Various pre-trained models are fine-tuned and used for this task. In this paper, we conduct a benchmark study of seven state-of-the-art pre-trained models for IQA of omnidirectional images. To this end, we first train these models using an omnidirectional database and compare their performance with the pre-trained versions. Then, we compare the use of viewports versus equirectangular (ERP) images as inputs to the models. Finally, for the viewports-based models, we explore the impact of the input number of viewports on the models’ performance. Experimental results demonstrated the performance gain of the re-trained CNNs compared to their pre-trained versions. Also, the viewports-based approach outperformed the ERP-based one independently of the number of selected views. Abderrezzaq Sendjasni, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh |
ICIP | 2 |
| 2020 | A Comprehensive Framework for 2D-JND Extension to 360-DEG ImagesabstractMasking effect is one of the most important perceptual properties that could be modeled by estimating an adaptive threshold known as the just noticeable difference (JND) referring to the maximum difference not perceived by the human visual system (HVS). In this paper, a novel framework to extend 2D-JND models to estimate thresholds for 360-degree images is proposed. The JND is estimated by viewports instead of applying it to the projected format image. Then, the viewport-based JND maps are back-projected to obtain the 360-JND map. To reduce the visible boundaries between viewports an alpha blending process is applied. The validation of the proposed framework is made using subjective experiments. It demonstrates that when applied to 360-degree images, the proposed framework outperforms the 2D-JND models in terms of observers preference at the same noise level. Sami Jaballah, Amegh Bhavsar, Mohamed-Chaker Larabi |
ICASSP | 3 |
| 2020 | Perceptual Versus Latitude-Based 360-Deg Video Coding OptimizationabstractIn this paper, we present two optimization strategies for the improvement of 360-degree video coding efficiency. For the first strategy, the just noticeable difference dedicated to 360-degree images is obtained thanks to a novel framework relying on rectilinear projections. The estimated JND is incorporated as a weighting function for the 360Lib-VTM-4.0 RDO scheme in order to make its selection results more correlated with the human perception. The second strategy consists of a QP adjustment according to the latitude position in the sphere, where the incorporated weighting factor are modeled using the Von Mises Distribution. In order to evaluate the performance of both video coding optimization strategies, a set of experimental tests are performed using four high resolution 360-degree video sequences. The RDO-based JND strategy outperforms the anchor 360Lib9.0-VTM 4.0 while preserving the perceived image quality. Further, the latitude-based strategy achieves higher performance when compared to 360Lib9.0-VTM 4.0 and to a state-of-the-art approach [1], in terms of BD-rate of WS-PSNR and CPP-PSNR. Sami Jaballah, Amegh Bhavsar, Mohamed-Chaker Larabi |
ICIP | 3 |
| 2020 | Eye Movement State Trajectory Estimator based on Ancestor SamplingabstractHuman gaze dynamics mainly concern about the sequence of the occurrence of three eye movements: fixations, saccades, and microsaccades. In this paper, we correlate them as three different states to velocities of eye movements. We build a state trajectory estimator based on ancestor sampling (ST EAS) model, which captures the features of the human temporal gaze pattern to identify the kind of visual stimuli. We used a gaze dataset of 72 viewers watching 60 video clips which are equally split into four visual categories. Uniformly sampled velocity vectors from the training set, are used to find the best suitable parameters of the proposed statistical model. Then, the optimized model is used for both gaze data classification and video retrieval on the test set. We observed 93.265% of classification accuracy and a mean reciprocal rank of 0.888 for video retrieval on the test set. Hence, this model can be used for viewer independent video indexing for providing viewers an easier way to navigate through the contents. Sai Phani Kumar Malladi, Jayanta Mukhopadhyay, Mohamed-Chaker Larabi, Santanu Chaudhury |
MMSP | 3 |
| 2019 | Just Noticeable Difference Model for Asymmetrically Distorted Stereoscopic ImagesabstractIn this paper, we propose a saliency-weighted stereoscopic JND (SSJND) model constructed based on psychophysical experiments, accounting for binocular disparity and spatial masking effects of the human visual system (HVS). Specifically, a disparity-aware binocular JND model is first developed using psychophysical data, and then is employed to estimate the JND threshold for non-occluded pixel (NOP). In addition, to derive a reliable 3D-JND prediction, we determine the visibility threshold for occluded pixel (OP) by including a robust 2D-JND model. Finally, SSJND thresholds of one view are obtained by weighting the resulting JND for NOP and OP with their visual saliency. Based on subjective experiments, we demonstrate that the proposed model outperforms the other 3D-JND models in terms of perceptual quality at the same noise level. Yu Fan 0001, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh, Christine Fernandez-Maloigne |
ICASSP | 2 |
| 2019 | Blind Stereopair Quality Assessment Using Statistics of Monocular and Binocular Image StructuresabstractIn this paper, we present a no-reference (NR) quality predictor for stereoscopic/3D images based on statistics aggregation of monocular and binocular local contrast features. In particular, for left and right views, we first extract statistical features of the image gradient magnitude (GM) and the Laplacian of Gaussian (LoG), describing the image local structures from different perspectives. The monocular statistical features are then combined to derive the binocular features based on a linear summation model using weightings based on LoG-response and image local-entropy, independently. These weights can effectively simulate the strength of the views dominance on binocular rivalry (BR) behavior of the human visual system. Subsequently, we further compute the GM features of the difference map between left and right views reflecting the distortion on disparity/depth information. Finally, the BR-inspired combined monocular and disparity-related binocular features associated with subjective quality scores are jointly used to construct a learned regression model relying on support vector machine regressor. Experimental results on three 3D-IQA benchmark databases demonstrate that our method achieves high quality prediction accuracy and competitive performance compared to state-of-the-art methods. Yu Fan 0001, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh |
ICIP | 2 |
| 2019 | Flexible Motion Vector Resolution Prediction for Video CodingabstractThe latest video coding standard, High Efficiency Video Coding (HEVC), uses quarter-pixel motion vector (MV) resolution for motion compensation. The adaptation of MV resolution supported by progressive MV resolution (PMVR) brings further improvement to performance by progressively adjusting the resolution according to the distance between the MV and its predictor. However, progressive adjustment of resolution by PMVR does not consider the inherent characteristics of the coding block. In this paper, we propose several ways to improve PMVR. First, we show that the performance of PMVR is correlated with the spatiotemporal characteristics of the video sequence. Then, to cope with the limitations of PMVR, we propose a flexible framework for the adaptation of MV resolution using: 1) PU size and gradient; 2) PU size, gradient, and MV components; and 3) PU size and spatiotemporal characteristics of the frames. Finally, a smart motion estimation around multiple MV predictors is performed to take full advantage of the proposed scheme. The proposed tools are implemented on top of HM-16.6. Extensive experiments and comparison with HEVC show 1.3%, 2.7%, and 1.0% average BD-Rate savings for random access, low-delay P, and low-delay B configurations, respectively. Bappaditya Ray, Mohamed-Chaker Larabi, Joël Jung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Asymmetric Dct-Jnd for Luminance Adaptation Effects: an Application To Perceptual Video Coding in Mv-HevcabstractBased on psychophysical experiments, we propose an asymmetric 3D just noticeable difference model (AJND) in the DCT domain taking into consideration the binocular properties of the human visual system (HVS), the background luminance and the spatial frequency of each DCT component. Subjective evaluations of the proposed AJND demonstrated that our proposed model offers good perceptual quality and could tolerate more distortion with PSNR reaching 24.19 dB. The proposed model has been used to perceptually optimize MV-HEVC by suppressing the residual transform coefficient that are lower than AJND values. Experimental results show that the proposed algorithm can achieve up to 14.62% bit-rate saving while preserving the perceived image quality. Sami Jaballah, Mohamed-Chaker Larabi, Jamel Belhadj Tahar |
ICASSP | 2 |
| 2018 | A Low-Complexity Video Encoder for Equirectangular Projected 360 Video Contentabstract360- video is gaining a lot of interest because of the immersive feeling brought by such a technology. Several projection formats are used to represent this type of content. Equirectangular projection (ERP) is one of the most widely used projection scheme for 360 panoramic content. The main drawback of ERP is its latitude dependent sampling density unlike conventional 2D content. Consequently, conventional 2D codecs such as HEVC are not optimal for the coding of ERP projected 360 content. To cope with this dependency, this work proposes an adaptation of motion vector resolution and minimum width of the coding block depending on its latitude. Experimental results show up to 0.5% BD-rate savings for motion contained sequences with 15% encoding time reduction in random access configuration. Bappaditya Ray, Joël Jung, Mohamed-Chaker Larabi |
ICASSP | 3 |
| 2018 | Towards Perceptually Guided Rate-Distortion Optimization For HevcabstractThis paper proposes a novel approach for perceptually guiding the rate-distortion optimization (RDO) process within the High Efficiency Video Coding (HEVC) standard. The reference codec does not consider effectively the perceptual characteristics of the input video and further, the particular perceptual sensitivity of each coding tree unit (CTU) inside a frame. The corresponding frame-level Lagrangian multiplier depends only on the quantization parameter. Inspired by the mechanisms of the human visual system, the proposed solution is a CTU-Ievel adjustment of the standard Lagrangian value based on a set of complementary measured features. These measures rely on the spatial and temporal analysis of the current CTU in the frequency domain. Based on perceptual quality indices and Bjontegaard delta measurements, over several resolutions of tested video sequences, the proposed method demonstrates a promising coding performance according to the rate-distortion compromise. Kais Rouis, Mohamed-Chaker Larabi, Jamel Belhadj Tahar |
ICASSP | 2 |
| 2018 | No-Reference Quality Assessment of Stereoscopic Images Based on Binocular Combination of Local Features StatisticsabstractNo-reference (NR) stereoscopic 3D (S3D) image quality assessment (SIQA) is still challenging due to the poor understanding of how the human visual system (HVS) judges image quality based on binocular vision. In this paper, we propose an efficient opinion-aware NR Stereoscopic Quality predictor based on local contrast statistics combination (SQSC). Specifically, for left and right views, we first extract statistical features of the gradient magnitude (GM) and Laplacian of Gaussian (LoG) responses, describing the image local structures from different perspectives. The HVS is insensitive to low-order statistical redundancies that can be removed by LoG filtering. Hence, the monocular statistical features are then fused to derive the binocular features based on a linear combination model using LoG responses-based weightings. These weightings can efficiently simulate the binocular rivalry (BR) phenomenon. Finally, the binocular features and the subjective scores were jointly employed to construct a learned regression model obtained by the support vector regression (SVR) algorithm. Experimental results on three widely used 3D IQA databases demonstrate the high prediction performance of the proposed method when compared to recent well performing SIQA methods. Yu Fan 0001, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh, Christine Fernandez-Maloigne |
ICIP | 2 |
| 2018 | Visual Attention for Rendered 3D ShapesabstractAbstract Understanding the attentional behavior of the human visual system when visualizing a rendered 3D shape is of great importance for many computer graphics applications. Eye tracking remains the only solution to explore this complex cognitive mechanism. Unfortunately, despite the large number of studies dedicated to images and videos, only a few eye tracking experiments have been conducted using 3D shapes. Thus, potential factors that may influence the human gaze in the specific setting of 3D rendering, are still to be understood. In this work, we conduct two eye‐tracking experiments involving 3D shapes, with both static and time‐varying camera positions. We propose a method for mapping eye fixations (i.e., where humans gaze) onto the 3D shapes with the aim to produce a benchmark of 3D meshes with fixation density maps, which is publicly available. First, the collected data is used to study the influence of shape, camera position, material and illumination on visual attention. We find that material and lighting have a significant influence on attention, as well as the camera path in the case of dynamic scenes. Then, we compare the performance of four representative state‐of‐the‐art mesh saliency models in predicting ground‐truth fixations using two different metrics. We show that, even combined with a center‐bias model, the performance of 3D saliency algorithms remains poor at predicting human fixations. To explain their weaknesses, we provide a qualitative analysis of the main factors that attract human attention. We finally provide a comparison of human‐eye fixations and Schelling points and show that their correlation is weak. Guillaume Lavoué, Frederic Cordier, Hyewon Seo, Mohamed-Chaker Larabi |
Comput. Graph. Forum | 4 |
| 2018 | Towards an automatic correction of over-exposure in photographs: Application to tone-mapping
Mekides Assefa Abebe, Alexandra Booth, Jonathan Kervec, Tania Pouli, Mohamed-Chaker Larabi |
Comput. Vis. Image Underst. | 5 |
| 2018 | Revertible tone mapping of high dynamic range imagery: Integration to JPEG 2000
Ines Bouzidi, Azza Ouled Zaid, Mohamed-Chaker Larabi |
Multim. Tools Appl. | 3 |
| 2018 | Low complexity intra prediction mode decision for 3D-HEVC depth coding
Sami Jaballah, Mohamed-Chaker Larabi, Jamel Belhadj Tahar |
Signal Process. Image Commun. | 2 |
| 2017 | Stereoscopic image quality assessment based on the binocular properties of the human visual systemabstractOne of the most challenging issues in stereoscopic image quality assessment (IQA) is how to effectively model the binocular behaviors of the human visual system (HVS). The latter has a great impact on the perceptual stereoscopic 3D (S3D) quality. This paper presents a stereoscopic IQA metric based on the properties of the HVS. Instead of measuring the quality of the left and the right views separately, the proposed method predicts the quality of a cyclopean image to ensure that the overall S3D quality is as close as possible to the binocular vision. The cyclopean image is synthesized based on the local entropy of each view with the aim to simulate the phenomena of the binocular rivalry/suppression. A 2D IQA metric is employed to assess the quality of both the cyclopean image and the disparity map. Additionally, the quality of the cyclopean image is modulated according to the visual importance of each pixel defined by the just noticeable difference (JND). Finally, the 3D quality score is derived by combining the quality estimates of the cyclopean image and disparity map. Experimental results show that the proposed method outperforms many other state-of-the-art SIQA methods in terms of prediction accuracy and computational efficiency. Yu Fan 0001, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh, Christine Fernandez-Maloigne |
ICASSP | 2 |
| 2017 | Full-reference stereoscopic image quality assessment accounting for binocular combination and disparity informationabstractOne of the most challenging issues in stereoscopic image quality assessment (SIQA) is how to effectively model the binocular behavior of the human visual system (HVS). The latter has a great impact on the perceptual 3D quality. In this paper, we propose a SIQA metric accounting for binocular combination properties and disparity information. Instead of computing the quality of the left and the right views separately, the proposed metric predicts the quality of a cyclopean image so as to have a good consistency with 3D human perception. The cyclopean image is synthesized based on the local entropy and the visual saliency of each view with the aim to simulate the phenomena of binocular fusion/rivalry. A 2D IQA metric is employed to assess the quality of both the cyclopean image and the disparity map. The obtained scores are used to derive the 3D quality score thanks to a pooling stage. Experimental results on three public 3D IQA databases show that the proposed method outperforms many other state-of-the-art SIQA methods, and achieves high prediction accuracy on these databases. Yu Fan 0001, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh, Christine Fernandez-Maloigne |
ICIP | 2 |
| 2017 | Blind image quality assessment in the complex frequency domainabstractIn this paper, we propose a no-reference (NR) image quality assessment (IQA) metric that operates in the complex frequency domain. A set of features are developed to model the natural scene statistics without depending on any specific visual distortion. The proposed approach relies on a statistical analysis of the transformed image, involving the importance of the phase and magnitude provided by the underlying complex coefficients. We further investigate the correlation between the different image spatial-frequency resolutions, i.e., representations under different scales and orientations in order to extract the directional features and energy distributions of an image. The validation of the NR metric is performed on a variety of challenging IQA databases and the obtained results show good correlation with subjective scores. Besides, the obtained performance is highly competitive compared to the top-performing NR IQA metrics. Kais Rouis, Mohamed-Chaker Larabi, Jamel Belhadj Tahar |
ICIP | 2 |
| 2017 | A block level adaptive MV resolution for video codingabstractThe latest video coding standard, HEVC, uses quarter pixel motion vector (MV) resolution for motion compensation. The adaptation of MV resolution supported by PMVR (progressive MV resolution) brings further improvement of the performance, by progressively adjusting the resolution according to the distance between the MV and the MV predictor (MVP). In this work, we propose to improve PMVR by adapting MV resolution at the prediction unit (PU) level relying on its size and its average absolute gradient. We additionally perform a smarter motion estimation around multiple MV predictors to fully take advantage of the proposed scheme. Compared to HEVC reference software (HM-16.6), the proposed method provides 1.2%, 3.2% and 1.2% average BD rate savings respectively for random access (RA), low-delay P (LP) and low-delay (LD) configurations. Bappaditya Ray, Joël Jung, Mohamed-Chaker Larabi |
ICME | 3 |
| 2017 | Toward an audiovisual attention model for multimodal video content
Naty Ould Sidaty, Mohamed-Chaker Larabi, Abdelhakim Saadane |
Neurocomputing | 2 |
| 2017 | Using distortion and asymmetry determination for blind stereoscopic image quality assessment strategy
Sid Ahmed Fezza, Aladine Chetouani, Mohamed-Chaker Larabi |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Perceptual Lightness Modeling for High-Dynamic-Range ImagingabstractThe human visual system (HVS) non-linearly processes light from the real world, allowing us to perceive detail over a wide range of illumination. Although models that describe this non-linearity are constructed based on psycho-visual experiments, they generally apply to a limited range of illumination and therefore may not fully explain the behavior of the HVS under more extreme illumination conditions. We propose a modified experimental protocol for measuring visual responses to emissive stimuli that do not require participant training, nor requiring the exclusion of non-expert participants. Furthermore, the protocol can be applied to stimuli covering an extended luminance range. Based on the outcome of our experiment, we propose a new model describing lightness response over an extended luminance range. The model can be integrated with existing color appearance models or perceptual color spaces. To demonstrate the effectiveness of our model in high dynamic range applications, we evaluate its suitability for dynamic range expansion relative to existing solutions. Mekides Assefa Abebe, Tania Pouli, Mohamed-Chaker Larabi, Erik Reinhard |
ACM Trans. Appl. Percept. | 3 |
| 2017 | Perceptually Driven Nonuniform Asymmetric Coding of Stereoscopic 3D VideoabstractAsymmetric stereoscopic video coding has already proven its effectiveness in reducing the bandwidth required for stereoscopic 3D delivery without degrading the visual quality. This approach, in which the left and right views are encoded with different levels of quality, relies on the perceptual theory of binocular suppression. However, to ensure comfortable 3D viewing, the just-noticeable level of asymmetry, i.e., the maximum quality gap between views, has to be carefully defined. Both subjectively and empirically fixed thresholds of asymmetry demonstrated either the maladjustment to content or dependency to the experimental design. This paper describes a new nonuniform asymmetric stereoscopic video coding method adaptively adjusting the level of asymmetry for each region of the image based on its perceptual significance. The proposed method uses a fully automated model that dynamically determines the best bounds of asymmetry for which the 3D viewing experience will not be altered. This is achieved by exploiting several human-visual-system-inspired models, namely, the binocular just-noticeable difference, and the visual saliency map and depth information. The simulation results show that the proposed method results in bit rate saving of up to 26% and provides better 3D visual quality compared with state-of-the-art asymmetric coding methods. Sid Ahmed Fezza, Mohamed-Chaker Larabi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | On the performance of 3D just noticeable difference modelsabstractThe just noticeable difference (JND) notion reflects the maximum tolerable distortion. It has been extensively used for the optimization of 2D applications. For stereoscopic 3D (S3D) content, this notion is different since it relies on different mechanisms linked to our binocular vision. Unlike 2D, 3D-JND models appeared recently and the related literature is rather limited. These models can be used for the sake of compression and quality assessment improvement for S3D content. In this paper, we propose a deep and comparative study of the existing 3D-JND models. Additionally, in order to analyze their performance, the 3D-JND models have been integrated in recent metric dedicated to stereoscopic image quality assessment (SIQA). The results are reported on two widely used S3D image databases. Yu Fan 0001, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh, Christine Fernandez-Maloigne |
ICIP | 2 |
| 2016 | Perceptually-adaptive quantization for stereoscopic video codingabstractIn this paper, we present a novel perceptually-based optimization for the improvement of stereoscopic video coding efficiency. The main idea of this proposed scheme is to adaptively adjust the quantization parameter by taking into account the Human Visual System perceptual characteristics. For this, a saliency map is generated from both views and then segmented into salient and non-salient regions. To make the proposed scheme effective, and inspired from the binocular suppression theory, the asymmetry is ensured by altering the saliency map and not the view. As a result, the proposed perceptual coding scheme effectively reduces the bit-budget without affecting the perceptual quality based on an optimization approach with asymmetric video coding taking into account the saliency map of each view. Experimental results on HEVC-MV show that the proposed algorithm can achieve over 20% bit-rate saving while preserving the perceived image quality. Sami Jaballah, Mohamed-Chaker Larabi, Jamel Belhadj Tahar |
ICIP | 2 |
| 2016 | On the Efficiency of Image Metrics for Evaluating the Visual Quality of 3D Modelsabstract3D meshes are deployed in a wide range of application processes (e.g., transmission, compression, simplification, watermarking and so on) which inevitably introduce geometric distortions that may alter the visual quality of the rendered data. Hence, efficient model-based perceptual metrics, operating on the geometry of the meshes being compared, have been recently introduced to control and predict these visual artifacts. However, since the 3D models are ultimately visualized on 2D screens, it seems legitimate to use images of the models (i.e., snapshots from different viewpoints) to evaluate their visual fidelity. In this work we investigate the use of image metrics to assess the visual quality of 3D models. For this goal, we conduct a wide-ranging study involving several 2D metrics, rendering algorithms, lighting conditions and pooling algorithms, as well as several mean opinion score databases. The collected data allow (1) to determine the best set of parameters to use for this image-based quality assessment approach and (2) to compare this approach to the best performing model-based metrics and determine for which use-case they are respectively adapted. We conclude by exploring several applications that illustrate the benefits of image-based quality assessment. Guillaume Lavoué, Mohamed-Chaker Larabi, Libor Vása |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | A visual attention model for stereoscopic 3D images using monocular cues
Iana Iatsun, Mohamed-Chaker Larabi, Christine Fernandez-Maloigne |
Signal Process. Image Commun. | 2 |
| 2015 | A case study in identifying acceptable bitrates for human face recognition tasks
Anastasia Tsifouti, Sophie Triantaphillidou, Mohamed-Chaker Larabi, Efthimia Bilissi, Alexandra Psarrou |
Signal Process. Image Commun. | 3 |
| 2014 | Asymmetric coding using Binocular Just Noticeable Difference and depth information for stereoscopic 3DabstractThe problem of determining the best level of asymmetry has been addressed by several recent works with the aim to guarantee an optimal binocular perception while keeping the minimum required information. To do so, subjective experiments have been conducted for the definition of an appropriate threshold. However, such an approach is lacking in terms of generalization because of the content variability. Moreover, using a fixed threshold does not allow an adaptation to the content and to the images' quality. The traditional asymmetric stereoscopic coding methods apply a uniform asymmetry by considering that all regions of an image have the same perceptual relevance which is not in compliance with the characteristics of human visual system (HVS). Consequently, this paper describes a fully automated model that dynamically determines the best bounds of asymmetry for each region of the image. Based on the Binocular Just Noticeable Difference (BJND) and the depth level in the scene, the proposed method achieves non-uniform reduction of spatial resolution of one view of the stereo pair with the aim to reduce bandwidth requirement. Experimental results show that the proposed method results in up to 43% of bitrate saving while outperforming the widely used asymmetric coding approaches in terms of 3D visual quality. Sid Ahmed Fezza, Mohamed-Chaker Larabi, Kamel Mohamed Faraoun |
ICASSP | 2 |
| 2014 | Using monocular depth cues for modeling stereoscopic 3D saliencyabstractSaliency is one of the most important features in human visual perception. It is widely used nowadays for perceptually optimizing image processing algorithms. Several models have been proposed for 2D images and only few attempts can be observed for 3D ones. In this paper, we propose a stereoscopic 3D saliency model relying on 2D saliency features jointly with depth obtained from monocular cues. On the one hand, the use of 2D saliency features is justified psychophysically by the similarity observed between 2D and 3D attention maps. On the other hand, 3D perception is significantly based on monocular cues. The validation of our model using state-of-the-art procedures including Kullback-Leibler divergence (KLD), area under the curve (AUC) and correlation coefficient (CC) in comparison with attention maps showed very good performance. Iana Iatsun, Mohamed-Chaker Larabi, Christine Fernandez-Maloigne |
ICASSP | 2 |
| 2014 | Stereoscopic image quality metric based on local entropy and binocular just noticeable differenceabstractDeveloping a metric that can reliably predict the perceptual 3D quality as perceived by the end user, is a challenging issue and a necessary tool for the success of 3D multimedia applications. The various attempts at predicting 3D quality of experience as the combination of 2D quality of the left and right images have shown their limitations, and particularly for the case of asymmetric distortions. In this paper we propose a full reference quality assessment metric for stereoscopic images based on the perceptual binocular characteristics. The proposed metric handles effectively the asymmetric distortions of stereoscopic images, by incorporating human visual system (HVS) characteristics. Our approach was motivated by the fact that in case of asymmetric quality, 3D perception mechanisms supports the view providing the most important and contrasted information. To achieve that, weighting factors are defined for the quality of each view according to the local information content. Add to that, to take into account the sensitivity of the HVS, quality score of each region are modulated based on the Binocular Just Noticeable Difference (BJND). Experimental results show that the proposed metric correlates better with human perception than the state-of-the-art metrics. Sid Ahmed Fezza, Mohamed-Chaker Larabi, Kamel Mohamed Faraoun |
ICIP | 2 |
| 2014 | Asymmetric coding of stereoscopic 3D based on perceptual significanceabstractAsymmetric stereoscopic coding is a very promising technique to decrease the bandwidth required for stereoscopic 3D delivery. However, one large obstacle is linked to the limit of asymmetric coding or the just noticeable threshold of asymmetry, so that 3D viewing experience is not altered. By way of subjective experiments, recent works have attempted to identify this asymmetry threshold. However, fixed threshold, highly dependent on the experiment design, do not allow to adapt to quality and content variation of the image. In this paper, we propose a new non-uniform asymmetric stereoscopic coding adjusting in a dynamic manner the level of asymmetry for each image region to ensure unaltered binocular perception. This is achieved by exploiting several HVS-inspired models; specifically we used the Binocular Just Noticeable Difference (BJND) combined with visual saliency map and depth information to quantify precisely the asymmetry threshold. Simulation results show that the proposed method results in up to 44% of bitrate saving and provides better 3D visual quality compared to state-of-the-art asymmetric coding methods. Sid Ahmed Fezza, Mohamed-Chaker Larabi, Kamel Mohamed Faraoun |
ICIP | 2 |
| 2014 | Spatio-temporal modeling of visual attention for stereoscopic 3D videoabstractModeling visual attention is an important stage for the optimization of image processing systems nowadays. Several models have been already developed for 2D static and dynamic content, but only few attempts can be found for stereoscopic 3D content. In this work we propose a saliency model for stereoscopic 3D video. This model is based the fusion of three maps i.e. spatial, temporal and depth. It relies on interest point features known for being close to human visual attention. Moreover, since 3D perception is mostly based on monocular cues, depth information is obtained using a monocular model predicting the depth position of objects. Several fusion strategies have been experimented in order to determine the best match for our model. Finally, our approach has been validated using state-of-the-art metrics in comparison to attention maps obtained by eye-tracking experiments, and showed good performance. Iana Iatsun, Mohamed-Chaker Larabi, Christine Fernandez-Maloigne |
ICIP | 2 |
| 2014 | A block artifact distortion measure for no reference video quality evaluationabstractIn This paper, we propose a new perceptually significant video quality metric to estimate the effect of block coding for standards H.264 AVC and MPEG2. Our method operates in the spatial domain and doesn't require a high complexity of computation. We compare it with Suthaharan's and MSU's techniques by using “LIVE” database. Results indicate that the proposed method outperforms Suthaharan's technique. They also indicate that our method is more effective than MSU's technique for the H.264 AVC and MPEG2 standard with the Spearman Rank Order Correlation Coefficient. Mohamed Ben Amor, Mohamed-Chaker Larabi, Fahmi Kammoun, Nouri Masmoudi |
IPAS | 2 |
| 2014 | Offline quality monitoring for legal evidence images in video-surveillance applications
Aldo Maalouf, Mohamed-Chaker Larabi, Didier Nicholson |
Multim. Tools Appl. | 2 |
| 2014 | A perceptual image completion approach based on a hierarchical optimization scheme
Trung Thanh Dang, Azeddine Beghdadi, Mohamed-Chaker Larabi |
Signal Process. | 3 |
| 2014 | Feature-Based Color Correction of Multiview Video for Coding and Rendering EnhancementabstractMultiview video (MVV) consists of capturing the same scene with multiple cameras from different viewpoints. Therefore, substantial illumination and color inconsistencies can be observed between different views. These color mismatches can significantly reduce compression efficiency and rendering quality. In this paper, we propose a preprocessing method for correcting these color discrepancies in MVV. To consider the occlusion problem, our method is based on an improvement of histogram matching (HM) algorithm using only common regions across views. These regions are defined by an invariant feature detector (scale invariant feature transform), followed by random sample consensus algorithm to increase the matching robustness. In addition, to maintain temporal correlation, HM algorithm is applied on a temporal sliding window, allowing to cope with time-varying acquiring system, camera moving capture, and real-time broadcasting. Moreover, unlike always choosing the center view as the reference by default, we propose an automatic selection algorithm based on both views statistics and quality. The experimental results show that the proposed method increases coding efficiency with gains of up to 1.1 and 2.2 dB for the luminance and chrominance components, respectively. Furthermore, once the correction is performed, the color of real and rendered views is harmonized and looks very consistent as a whole. Sid Ahmed Fezza, Mohamed-Chaker Larabi, Kamel Mohamed Faraoun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Perceptual quality assessment for color image inpaintingabstractA novel objective measure for assessing the quality of image in-painting is proposed. In contrast to standard image quality metrics, the proposed one takes into account some constraints and characteristics related to the specific goals of inpainting techniques. The idea is to combine spatial low-level features and perceptual criteria in the design of the objective Image Inpainting Quality Metric (IIQM). The used characteristics are the visual coherence of the recovered regions and the visual saliency describing the visual importance of an area. Experimental results demonstrate the good performance of the proposed IIQM and its well adaptation to the evaluation of image inpainting results. Thanh Trung Dang, Azeddine Beghdadi, Mohamed-Chaker Larabi |
ICIP | 3 |
| 2013 | Perceptual Metrics for Static and Dynamic Triangle MeshesabstractAbstract Almost all mesh processing procedures cause some more or less visible changes in the appearance of objects represented by polygonal meshes. In many cases, such as mesh watermarking, simplification or lossy compression, the objective is to make the change in appearance negligible, or as small as possible, given some other constraints. Measuring the amount of distortion requires taking into account the final purpose of the data. In many applications, the final consumer of the data is a human observer, and therefore the perceptibility of the introduced appearance change by a human observer should be the criterion that is taken into account when designing and configuring the processing algorithms. In this review, we discuss the existing comparison metrics for static and dynamic (animated) triangle meshes. We describe the concepts used in perception‐oriented metrics used for 2D image comparison, and we show how these concepts are employed in existing 3D mesh metrics. We describe the character of subjective data used for evaluation of mesh metrics and provide comparison results identifying the advantages and drawbacks of each method. Finally, we also discuss employing the perception‐correlated metrics in perception‐oriented mesh processing algorithms. Massimiliano Corsini, Mohamed-Chaker Larabi, Guillaume Lavoué, Oldrich Petrík, Libor Vása, Kai Wang 0002 |
Comput. Graph. Forum | 2 |
| 2013 | Biologically inspired approaches for visual information processing and analysis
Azeddine Beghdadi, Abdesselam Bouzerdoum, Khan M. Iftekharuddin, Mohamed-Chaker Larabi |
Signal Process. Image Commun. | 4 |
| 2013 | A survey of perceptual image processing methods
Azeddine Beghdadi, Mohamed-Chaker Larabi, Abdesselam Bouzerdoum, Khan M. Iftekharuddin |
Signal Process. Image Commun. | 2 |
| 2013 | Attentional mechanisms driven adaptive quantization and selective bit allocation scheme for H.264/AVC
Miryem Hrarti, Abdelhakim Saadane, Mohamed-Chaker Larabi, Rémi Barland |
Signal Process. Image Commun. | 3 |
| 2012 | Robust foveal wavelet-based object trackingabstractIn this work, a foveal wavelet-based Mean Shift Tracking Algorithm is presented. The foveal wavelets introduced by Mallat [16] are known by their high capability to precisely characterize the holder regularity of singularities. Therefore, by using the foveal wavelet transform, image features are accurately identified and are well discriminated from noise. These wavelets are used to extract the texture features of the target object. The extracted features are then used to construct a joint color-foveal textures histogram to represent the target object. Once the joint histogram is obtained, it is applied to the mean shift framework in order to track a target object in a video sequence. The experimental results showed that the proposed approach overcomes the traditional mean shift tracking technique as well as other existing tracking algorithms. Aldo Maalouf, Mohamed-Chaker Larabi |
ICASSP | 2 |
| 2012 | An efficient demosaicing technique using geometrical informationabstractColor image sensors use color filter arrays (CFA) to capture information at each sensor pixel position and require color demosaicing to reconstruct full color images. The quality of the demosaicked image is hindered by the sensor characteristics during the acquisition process. In this work, we propose a bandelet-based demosaicing method for color images. To this end, we have used a spatial multiplexing model of color in order to obtain the luminance and the chrominance components of the acquired image. Then, a luminance filter is used to reconstruct the luminance component. Thereafter, based on the concept of maximal gradient of multivalued images, we propose an extension of the bandelet representation for the case of multivalued images. Finally, demosaicing is performed by merging the luminance and each of the chrominance component in the multivalued bandelet transform domain. The experimental evaluation of the proposed scheme shows beneficial performance over existing demosaicing approaches. Aldo Maalouf, Mohamed-Chaker Larabi, Sabine Süsstrunk |
ICIP | 2 |
| 2011 | CYCLOP: A stereo color image quality assessment metricabstractIn this work, a reduced reference (RR) perceptual quality metric for color stereoscopic images is presented. Given a reference stereo pair of images and their "distorted" version, we first compute the disparity map of both the reference and the distorted stereoscopic images. To this end, we define a method for color image disparity estimation based on the structure tensors properties and eigenvalues/eigenvectors analysis. Then, we compute the cyclopean images of both the reference and the distorted pairs. Thereafter, we apply a multispectral wavelet decomposition to the two cyclopean color images in order to describe the different channels in the human visual system (HVS). Then, contrast sensitivity function (CSF) filtering is performed to obtain the same visual sensitivity information within the original and the distorted cyclopean images. Thereafter, based on the properties of the human visual system (HVS), rational sensitivity thresholding is performed to obtain the sensitivity coefficients of the cyclopean images. Finally, RR stereo color image quality assessment (SCIQA) is performed by comparing the sensitivity coefficients of the cyclopean images and studying the coherence between the disparity maps of the reference and the distorted pairs. Experiments performed on color stereoscopic images indicate that the objective scores obtained by the proposed metric agree well with the subjective assessment scores. Aldo Maalouf, Mohamed-Chaker Larabi |
ICASSP | 2 |
| 2011 | A robust content-based JPWL transmission over a realistic MIMO channel under perceptual constraintsabstractThis paper proposes a global approach of JPWL (ISO/IEC 15444-11) image transmission over a realistic wireless channel able to ensure the best Quality of Service (QoS). In order to exploit the channel diversity, we consider a Closed-Loop MIMO-OFDM scheme with different precoder designs. In particular, the high flexibility of QoS precoder allows taking into account the scalability of JPWL jointly with the instantaneous MIMO channel status. This increases the visual quality of received images. The monitoring of the quality is made by a reduced-reference metric (QIP) based on object's saliency and interest point, both linked to human perception. It is performed in association with a robust JPWL decoder to determine the optimal decoding configuration in terms of PSNR. The proposed scheme provides very good results and its performance is shown through a realistic wireless channel. Julien Abot, Michael Nauge 0001, Clency Perrine, Mohamed-Chaker Larabi, Cyril Bergeron, Yannis Pousset, Christian Olivier |
ICIP | 4 |
| 2010 | Bandelet-based stereo image codingabstractIn this work, a bandelet-based coding scheme for stereo images is presented. The proposed scheme efficiently integrates the coding of the disparity map with the reference image. The disparity map is obtained via disparity estimation in a geometric similarity framework. The scheme first computes the bandelet transform of both left and right images. Consequently, each image is segmented into a quadtree where each dyadic square regroups pixels sharing the same geometric flow direction. Then, the disparity map is obtained by studying the geometric similarities between the dyadic squares of both images quadtrees. This is accomplished by the minimization of a cost measure function that is defined on the geometric properties of the quadtrees. Finally, the bandelet transform coefficients of the reference and residual images with the disparity map are encoded and transmitted in partitions (squares) which leads to lower entropy. The experimental evaluation of the proposed scheme shows beneficial performance over other stereoscopic coders in the literature. Aldo Maalouf, Mohamed-Chaker Larabi |
ICASSP | 2 |
| 2010 | Image retargeting using a bandelet-based similarity measureabstractMedia content retargeting aims to adapt images/videos to displays of large or small sizes. In this work, we propose a bandelet-based image retargeting algorithm for summarizing image data into smaller sizes. First, we define a multi-scale bandelet-based perceptual similarity measure which measures the geometric and perceptual similarities between two images at different bandelet scales. Two images are said to be geometrically similar if they have approximately the same geometric flow and quadtree structure. After determining the geometric similarity, a perceptual similarity measure based on the properties of the human visual system is defined to assess the perceptual difference between the original image and the retargeted one. Then, the problem of image retargeting is considered as a geometric optimization problem based on the bandelet-based geometric and perceptual similarity measures. That is, for an image S we search for a retargeted image T that contains as much as possible of geometric and perceptual information from S and, consequently, preserves visual coherence. The proposed retargeting algorithm outperforms the state-of-the-art methods in terms of the visual quality of the retargeted image. Aldo Maalouf, Mohamed-Chaker Larabi |
ICASSP | 2 |
| 2010 | Stereo image coding based on binocular energy modelingabstractStereoscopic imaging technologies are seen as the next generation of visual presentation, improving the quality of experience of the viewer. It uses two different sequences acquired from two regular cameras or from a regular camera with an additional specific depth camera. This means that the size of data is at least doubled. Thus, the coding process becomes very crucial. In this framework, we propose a stereoscopic coder based on visual properties. The matching of two images is computed by a binocular energy model based on the simple and complex cells functions allowing the fusion of both retinal images in the visual cortex. Mathematical functions were used to reproduce the behavior of these cells particularly complex wavelet transform (CWT) and bandelet transform. Our coder output is a disparity map, a residual image and the reference image. The innovative part of this work lies in a matching technique based on the binocular energy. The results are presented in comparative curves with one of the most known coder in literature. Rafik Bensalma, Mohamed-Chaker Larabi |
ICIP | 2 |
| 2010 | Towards a perceptual quality metric for color stereo imagesabstractIn this paper, we propose a quality metric for color stereo images. The concept of our metric is inspired by the behavior of simple and complex cells located in the primary visual cortex. These cells are responsible for merging left and right retinal images. To replicate the task performed by these cells, we adopted an approach based on spatial-frequency transform with the processing of selective orientations. From that, a model that calculates the binocular energy contained in the left and right retinal images has been proposed. The amplitude variation of the binocular energy defines the quality criterion of the reconstructed depth within the Human Visual System (HVS). Finally, from the experimental results, the used criterion seems to be correlated to human judgment obtained by psychophysical tests. Rafik Bensalma, Mohamed-Chaker Larabi |
ICIP | 2 |
| 2010 | A reduced-reference metric based on the interest points in color imagesabstractIn the last decade, an important research effort has been dedicated to quality assessment from subjective and objective points of view. The focus was mainly on Full Reference (FR) metrics because of the ability to compare to an original. Only few works were oriented to Reduced Reference (RR) or No Reference (NR) metrics, very useful for applications where the original image is not available such as transmission or monitoring. In this work, we propose a RR metric based on two concepts, the interest points of the image and the objects saliency on color images. This metric needs a very low amount of data (lower than 8 bytes) to be able to compute the quality scores. The results show a high correlation between the metric scores and the human judgement and a better quality range than well-known metrics like PSNR or SSIM. Finally, interest points have shown that they can predict the quality of compressed color images. Michael Nauge 0001, Mohamed-Chaker Larabi, Christine Fernandez-Maloigne |
PCS | 2 |
| 2009 | Camera motion influence on dynamic saliency central biasabstractSaliency models have been extensively studied for static images and the focus is now on moving images. There is a central bias in both cases that is emphasized in the dynamic case. One aspect in this latter is the camera motion that influences the scene interpretation. The movie director exploits this motion to make the observer focus on the targeted object which is often in the center of the scene. This aspect is not taken into account in current saliency dynamic models. In this paper, we study the camera motion influence on the gaze distribution in order to include it in a new saliency model. Observers' gazes are recorded with an eye tracker, camera motions (e.g. tracking, zoom ...) are calculated thanks to a polynomial projection of the motion field and the motion influence is statistically tested on the recorded gazes. Étienne Baudrier, Vincent Rosselli, Mohamed-Chaker Larabi |
ICASSP | 3 |
| 2009 | New H.264 intra-rate estimation and inter-rate control driven by improved MAD-based Contrast SensitivityabstractThis paper aims to improve H.264 bit-rate control. The proposed algorithm is based on a new and efficient rate-quantization (R-Q) model for the intra frame. For the inter frame, we propose to replace the current use of MAD by a new MAD-based human contrast sensitivity (MAD-CS) which is a more accurate complexity measure. R-Q model for the intra frame results from extensive experiments. The optimal initial quantization parameter QP is based on both target bit-rate and complexity of I-frame. The I-frame target bit-rate is derived from the global target bit-rate by using a new non linear model. MAD-CS includes the contrast sensitivity of the human visual system and weights the absolute differences by the probability of their occurrence. Extensive simulation results show that the use of MAD-CS and the proposed R-Q model achieves better rate control for intra frames, reduces the bit-rates when compared to the H.264 rate control adopted in JM reference software, minimizes the peak to signal ratio variations among encoded pictures and increases significantly as well subjective visual quality (measured by psycho visual experiments) as objective one. Miryem Hrarti, Hakim Saadane, Mohamed-Chaker Larabi, Ahmed Tamtaoui, Driss Aboutajdine |
ICIP | 3 |
| 2009 | Low-complexity enhanced lapped transform for image coding in JPEG XR / HD photoabstractJPEG-XR is a new image compression standard that aims at achieving state-of-the-art image compression, while simultaneously keeping the encoder and decoder complexities as low as possible. JPEG-XR is based on Microsoft technology known as HDPHOTO and makes use of a block-transform. This transform, known as Lapped Biorthogonal Transform (LBT), requires only a small memory footprint while providing the compression benefits of a larger block transform. In this work, we propose to replace the LBT by a representation in Legendre orthogonal polynomial basis. The motivation behind using the Legendre polynomials is that, in general, moment functions of orthogonal polynomials provide better feature representations over other type of moments and have some properties related to the human visual system (HVS). However, Legendre polynomials have a unit weight function and recurrence relation involving real coefficients, which make them suitable for defining image representation. We show that the expansion in Legendre polynomial basis can be implemented via lifting operations and has the same computation complexity as the LBT. The experimental evaluation of our modified JPEG-XR scheme shows beneficial improvements in terms of visual quality over the standard JPEG-XR. Aldo Maalouf, Mohamed-Chaker Larabi |
ICIP | 2 |
| 2009 | Still image coding using a wavelet-like transformabstractIn this paper, a new image coding scheme based on a wavelet-like transform derived from orthogonal polynomial basis is presented. From a set of bivariate orthogonal polynomial functions, we first obtain the 2D non-separable wavelet functions to propose a wavelet-like transform coding. The motivation behind using orthogonal polynomials is that they exhibit some properties related to the human visual system (HVS). After applying the proposed transformation, the transform coefficients are threshold coded using quantization and bit allocation as in JPEG baseline system. The performance of the proposed transform coding is reported. The proposed coding scheme is also compared with other transform coding methods such as JPEG, JPEG 2000 and JPEG-XR/HDPHOTO. Aldo Maalouf, Mohamed-Chaker Larabi, Christine Fernandez-Maloigne |
ICIP | 2 |
| 2008 | Subjective and Objective Assesment of Visual Image Quality Metrics and Still Image CodecsabstractSummary form only given. Objective quality assessment of lossy image compression codecs have become an important part of the recent call of the JPEG committee for advanced image coding. We evaluated JPEG with Huffman and arithmetic coding option, a visual and PSNR optimal JPEG2000 version, H.264/AVC and the recently proposed HDPhoto format by Microsoft. For objective evaluation, we use a color version of the M-SSIM metric and the high-dynamic range version of VDP. The results obtained from these tests are compared to subjective testing obtained from an ordering test run by 15 observers in two sessions. Subjective results are compiled to Mean Opinion Score (MOS) and passed through a Kurtosis test to verify their validity and to reject outliers. Thomas Richter 0005, Mohamed-Chaker Larabi |
DCC | 2 |
| 2008 | Image Rendering Based on a Spatial Extension of the CIECAM02abstractWith the multiplicity of imaging devices, the color quality and portability have become a very challenging problem. Moreover, a color is perceived with regards to its environment. So, if this environment changes it implies a change in the perceived color. In order to address this influence, the CIE (Commission Internationale de I'eclairage) has standardized a tool named color appearance model (CIECAM97*, CIECAM02). These models are able to take into account many phenomena related to human vision of color and can predict the color of a stimulus, function of its observations conditions. However, these models do not deal with the influence of spatial frequencies which can have a big impact on our perception. In this paper, we present an extended version of the CIECAM02 that integrates a spatial model correcting the color in relation to its spatial frequency. Moreover, the previous model has been modified to deal with images and not only single stimulus. The main difference with the rendering models (e.g. iCAM) lies in the fact that the proposed model, takes into account the spatial repartition of a pixel in addition to its environment. The obtained results are sound and demonstrate the efficiency of the proposed extension. This has been checked thanks to a psychophysical study where observers were assigned the task of assessing the quality of the improved version in comparison to the original. Olivier Tulet, Mohamed-Chaker Larabi, Christine Fernandez-Maloigne |
WACV | 2 |
| 2006 | A Novel Approach for Constructing an Achromatic Contrast Sensitivity Function by MatchingabstractModels of the human visual system are particularly interesting to quantify the quality of the display systems and to predict if visual information will be perceptible or not. One of these models is the contrast sensitivity function (CSF) which characterizes the sensitivity of the visual system to the spatial and temporal frequencies. The achromatic CSF can be measured, relatively, by a method of pairing which consists in matching the contrast of a test grid with that of a reference grid. To determine the reproducible grids on a screen, it is practical to use a frequency/observation distance diagram. The tests of this study are carried out under the conditions of medical diagnosis for radiographies with sinusoidal stimuli. The obtained results were approximated by a model in order to facilitate their integration in other models. Mohamed-Chaker Larabi, Vincent Brodbeck, Christine Fernandez-Maloigne |
ICIP | 1 |