EDBT 2026 Demo / reviewers in the wild / expert
Sid Ahmed Fezza
dblp:15/9811
· DBLP profile ↗
29ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0001-6453-8588ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 9 first-author · 14 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos: Methods and Results
Yixu Chen, Hai Wei, Pierre R. Lebreton, Patrick Le Callet, Alexander Kopte, Amritha Premkumar, Anna Meyer, Baojun Li, Changsheng Gao, Christian Herglotz, Christian Timmerer, Dandan Zhu 0001, Diwakara Reddy, Dong Liu 0002, Dounia Hammou, Guangtao Zhai, Hadi Amirpour, Hao Cheng 0015, Hichem Faraoun, Jonas Janzen, Krishna Srikar Durbha, Li Li 0040, Marc Windsheimer, MohammadAli Hamidi, Mykyta Skipenko, Paul Wawerek-Lopez, Pragyadipta Adhya, Prajit T. Rajendran, Rafal Mantiuk, Shien Ke, Sid Ahmed Fezza, Simon Deniffel, Wei Sun 0029, Weixia Zhang, Xiangguang Chen, Zuowei Cao, Minhao Tang, Xiaoyan Sun 0001, Xingwei Liu, Yeganeh Chatri, Yenan Xu |
QoMEX | 32 |
| 2026 | Complexity prediction of hardware and software video transcoding in the cloud
Taieb Chachou, Sid Ahmed Fezza, Wassim Hamidouche, Ghalem Belalem, Hadi Amirpour |
Multim. Tools Appl. | 2 |
| 2026 | Does data augmentation help or hinder the generalization of deepfake video detection?
Bachir Kaddar, Sid Ahmed Fezza, Elhocine Boutellaa, Wassim Hamidouche, Abdenour Hadid |
Multim. Tools Appl. | 2 |
| 2025 | Energy Backdoor Attack to Deep Neural NetworksabstractThe rise of deep learning (DL) has increased computing complexity and energy use, prompting the adoption of application specific integrated circuits (ASICs) for energy-efficient edge and mobile deployment. However, recent studies have demonstrated the vulnerability of these accelerators to energy attacks. Despite the development of various inference time energy attacks in prior research, backdoor energy attacks remain unexplored. In this paper, we design an innovative energy backdoor attack against deep neural networks (DNNs) operating on sparsity-based accelerators. Our attack is carried out in two distinct phases: backdoor injection and backdoor stealthiness. Experimental results using ResNet-18 and MobileNet-V2 models trained on CIFAR-10 and Tiny ImageNet datasets show the effectiveness of our proposed attack in increasing energy consumption on trigger samples while preserving the model’s performance for clean/regular inputs. This demonstrates the vulnerability of DNNs to energy backdoor attacks. The source code of our attack is available at: https://github.com/hbrachemi/energybackdoor. Hanene Brachemi Meftah, Wassim Hamidouche, Sid Ahmed Fezza, Olivier Déforges, Kassem Kallas |
ICASSP | 3 |
| 2025 | Convex Hull Prediction Methods for Bitrate Ladder Construction: Design, Evaluation, and ComparisonabstractHTTP adaptive streaming (HAS) has emerged as a prevalent approach for over-the-top (OTT) video streaming services due to its ability to deliver a seamless user experience. A fundamental component of HAS is the bitrate ladder, which comprises a set of encoding parameters (e.g., bitrate-resolution pairs) used to encode the source video into multiple representations. This adaptive bitrate ladder enables the client’s video player to dynamically adjust the quality of the video stream in real-time based on fluctuations in network conditions, ensuring uninterrupted playback by selecting the most suitable representation for the available bandwidth. The most straightforward approach involves using a fixed bitrate ladder for all videos, consisting of pre-determined bitrate-resolution pairs known as one-size-fits-all . Conversely, the most reliable technique relies on intensively encoding all resolutions over a wide range of bitrates to build the convex hull , thereby optimizing the bitrate ladder by selecting the representations from the convex hull for each specific video. Several techniques have been proposed to predict content-based ladders without performing a costly, exhaustive search encoding. This article provides a comprehensive review of various convex hull prediction methods, including both conventional and learning-based approaches. Furthermore, we conduct a benchmark study of several handcrafted- and deep learning (DL)-based approaches for predicting content-optimized convex hulls across multiple codec settings. The considered methods are evaluated on our proposed large-scale dataset, which includes 300 UHD video shots encoded with software and hardware encoders using three state-of-the-art video standards, including AVC/H.264, HEVC/H.265, and VVC/H.266, at various bitrate points. Our analysis provides valuable insights and establishes baseline performance for future research in this field ( Dataset URL : https://nasext-vaader.insa-rennes.fr/ietr-vaader/datasets/br_ladder ). Ahmed Telili, Wassim Hamidouche, Hadi Amirpour, Sid Ahmed Fezza, Christian Timmerer, Luce Morin |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Deepfake Detection Using Spatiotemporal TransformerabstractRecent advances in generative models and the availability of large-scale benchmarks have made deepfake video generation and manipulation easier. Nowadays, the number of new hyper-realistic deepfake videos used for negative purposes is dramatically increasing, thus creating the need for effective deepfake detection methods. Although many existing deepfake detection approaches, particularly CNN-based methods, show promising results, they suffer from several drawbacks. In general, poor generalization results have been obtained under unseen/new deepfake generation methods. The crucial reason for the above defect is that CNN-based methods focus on the local spatial artifacts, which are unique for every manipulation method. Therefore, it is hard to learn the general forgery traces of different manipulation methods without considering the dependencies that extend beyond the local receptive field. To address this problem, this article proposes a framework that combines Convolutional Neural Network (CNN) with Vision Transformer (ViT) to improve detection accuracy and enhance generalizability. Our method, namedHCiT, exploits the advantages of CNNs to extract meaningful local features, as well as the ViT’s self-attention mechanism to learn discriminative global contextual dependencies in a frame-level image explicitly. In this hybrid architecture, the high-level feature maps extracted from the CNN are fed into the ViT model that determines whether a specific video is fake or real. Experiments were performed on Faceforensics++, DeepFake Detection Challenge preview, Celeb datasets, and the results show that the proposed method significantly outperforms the state-of-the-art methods. In addition, the HCiT method shows a great capacity for generalization on datasets covering various techniques of deepfake generation. The source code is available at: https://github.com/KADDAR-Bachir/HCiT Bachir Kaddar, Sid Ahmed Fezza, Zahid Akhtar, Wassim Hamidouche, Abdenour Hadid, Joan Serra-Sagristà |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | 2BiVQA: Double Bi-LSTM-based Video Quality Assessment of UGC VideosabstractRecently, with the growing popularity of mobile devices as well as video sharing platforms (e.g., YouTube, Facebook, TikTok, and Twitch), User-Generated Content (UGC) videos have become increasingly common and now account for a large portion of multimedia traffic on the internet. Unlike professionally generated videos produced by filmmakers and videographers, typically, UGC videos contain multiple authentic distortions, generally introduced during capture and processing by naive users. Quality prediction of UGC videos is of paramount importance to optimize and monitor their processing in hosting platforms, such as their coding, transcoding, and streaming. However, blind quality prediction of UGC is quite challenging, because the degradations of UGC videos are unknown and very diverse, in addition to the unavailability of pristine reference. Therefore, in this article, we propose an accurate and efficient Blind Video Quality Assessment (BVQA) model for UGC videos, which we name 2BiVQA for double Bi-LSTM Video Quality Assessment. 2BiVQA metric consists of three main blocks, including a pre-trained Convolutional Neural Network to extract discriminative features from image patches, which are then fed into two Recurrent Neural Networks for spatial and temporal pooling. Specifically, we use two Bi-directional Long Short-term Memory networks, the first is used to capture short-range dependencies between image patches, while the second allows capturing long-range dependencies between frames to account for the temporal memory effect. Experimental results on recent large-scale UGC VQA datasets show that 2BiVQA achieves high performance at lower computational cost than most state-of-the-art VQA models. The source code of our 2BiVQA metric is made publicly available at https://github.com/atelili/2BiVQA . Ahmed Telili, Sid Ahmed Fezza, Wassim Hamidouche, Hanene Brachemi Meftah |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Efficient Per-Shot Transformer-Based Bitrate Ladder Prediction for Adaptive Video StreamingabstractRecently, HTTP adaptive streaming (HAS) has become a standard approach for over-the-top (OTT)-based video streaming services due to its ability to provide smooth streaming. In HAS, stream representations are encoded to target a specific bitrate providing a wide range of operating bitrates known as the bitrate ladder. In the past, a fixed bitrate ladder approach for all videos has been widely used. However, such a method does not consider video content, which can vary considerably in motion, texture, and scene complexity. Moreover, building a per-title bitrate ladder based on an exhaustive encoding is quite expensive due to the large encoding parameter space. Thus, alternative solutions allowing accurate and efficient per-title bitrate ladder prediction are in great demand. On the other hand, self-attention-based architectures have achieved tremendous performance in large language models (LLMs) and particularly vision transformers (ViTs) in computer vision tasks. Therefore, this paper investigates ViT’s capabilities in building an efficient bitrate ladder without performing any encoding process. We provide the first in-depth analysis of the prediction accuracy and the complexity overhead induced by the ViTs model in predicting the bitrate ladder on a large and diverse video dataset. The source code of the proposed solution and the dataset will be made publicly available. Ahmed Telili, Wassim Hamidouche, Sid Ahmed Fezza, Luce Morin |
ICIP | 3 |
| 2023 | Energy Consumption and Carbon Footprint of Modern Video Decoding SoftwareabstractThe estimation of energy consumption has become vital in developing eco-friendly and sustainable video streaming solutions to monitor CO2 emissions. In this paper, we seek to evaluate and compare the energy consumption and CO2 emissions of the decoding process related to three popular video coding standards, namely AVC, HEVC, VVC, along with two video formats VP9, and AV1 through their real-time software decoders, including h264, hevc, VVdeC/OpenVVC, vp9, and libdav1d. The evaluation is conducted on two types of consumer hardware, desktop PC and laptop. To ensure a fair evaluation, we also assess the coding efficiency of software encoder implementations using three objective quality metrics. The experimental results revealed that the h264 decoder consumes the lowest energy and is associated with the lowest CO2 emissions compared to other decoders on both hardware platforms. On the other hand, the VVenC encoder enhances coding efficiency at the cost of increased decoding energy consumption and CO2 emissions, particularly noticeable in the case of the OpenVVC decoder. Meanwhile, x265/hevc achieves a compelling balance between coding efficiency and decoding energy consumption. The full results of this work are available at https://decodingenergy.github.io/decoding_energy_co2.html. Taieb Chachou, Wassim Hamidouche, Sid Ahmed Fezza, Ghalem Belalem |
MMSP | 3 |
| 2022 | Evaluation of Pre-Trained CNN Models for Geographic Fake Image DetectionabstractThanks to the remarkable advances in generative adversarial networks (GANs), it is becoming increasingly easy to generate/manipulate images. The existing works have mainly focused on deepfake in face images and videos. However, we are currently witnessing the emergence of fake satellite images, which can be misleading or even threatening to national security. Consequently, there is an urgent need to develop detection methods capable of distinguishing between real and fake satellite images. To advance the field, in this paper, we explore the suitability of several convolutional neural network (CNN) architectures for fake satellite image detection. Specifically, we benchmark four CNN models by conducting extensive experiments to evaluate their performance and robustness against various image distortions. This work allows the establishment of new baselines and may be useful for the development of CNN-based methods for fake satellite image detection. Sid Ahmed Fezza, Mohammed Yasser Ouis, Bachir Kaddar, Wassim Hamidouche, Abdenour Hadid |
MMSP | 1 |
| 2022 | Visual Security Evaluation of Perceptually Encrypted Images based on Multi-Task LearningabstractOver past decades, many image encryption algorithms have been proposed, among which we can cite the perceptual/selective encryption methods which have attracted wide attention. Such methods allow for adjusting the scrambling intensity, it is therefore essential to have a reliable visual security metric to adjust the scrambling intensity on the one hand and to evaluate the visual security of encrypted images on the other hand. Usually, these tasks are performed based on classical randomness-based measures or image quality assessment metrics. However, these methods have shown their inadequacy as a visual security metric, as they do not address content intelligibility, which represents an essential security requirement. Moreover, these methods are either dedicated to the prediction of visual security (VS) or visual quality (VQ), but not both. In this paper, we propose a no-reference (NR) visual security metric for perceptually encrypted images based on deep multi-task learning, which we dub the Multi-Task Visual Security (MTVS) metric. The proposed metric consists of one shared convolutional neural network (CNN) followed by two separate sub-networks of fully-connected (FC) layers, where one sub-network is responsible for predicting the VS score, while the other is for predicting the VQ score. Experiments were performed on two publicly perceptually encrypted image databases and the results show that the proposed metric yields superior performance on both VS and VQ prediction tasks. The source code and models are available at: https://github.com/Mamadou-Keita/MTVS. Mamadou Keita, Sid Ahmed Fezza, Wassim Hamidouche, Azeddine Beghdadi |
MMSP | 2 |
| 2022 | Benchmarking Learning-based Bitrate Ladder Prediction Methods for Adaptive Video StreamingabstractHTTP adaptive streaming (HAS) is increasingly adopted by over-the-top (OTT)-based video streaming services, it allows clients to dynamically switch among various stream representations. Each of these representations is encoded to target a specific bitrate providing a wide range of operating bitrates known as the bitrate ladder. Several approaches with different levels of complexity are currently used to build such a bitrate ladder. The most straightforward method is to use a fixed bitrate ladder for all videos, which is a set of bitrate-resolution pairs, called “one-size-fits-all”, and the most complex is based on the intensive encoding of all resolutions over a wide bitrate range to construct the convex-hull. This latter is then used to obtain a per-title bitrate ladder. Recently, various methods relying on machine learning (ML) techniques have been proposed to predict content-based ladder without performing exhaustive search encoding. In this paper, we conduct a benchmark study of several handcrafted and deep learning (DL)-based approaches for predicting content-optimized bitrate ladder, which we believe provides baseline methods and will be useful for future research in this field. The obtained results, based on 200 video sequences compressed with the high-efficiency video coding (HEVC) encoder, reveal that the most efficient method predicts the bitrate ladder without performing any encoding process at the cost of a slight Bjøntegaard delta bitrate (BD-BR) loss of 1.43% compared to the exhaustive approach. The dataset and the source code of the considered methods are made publicly available at: https://github.com/atelili/Bitrate-Ladder-Benchmark. Ahmed Telili, Wassim Hamidouche, Sid Ahmed Fezza, Luce Morin |
PCS | 3 |
| 2022 | Deep multi-task learning for image/video distortions identification
Zoubida Ameur, Sid Ahmed Fezza, Wassim Hamidouche |
Neural Comput. Appl. | 2 |
| 2022 | Detect and defense against adversarial examples in deep learning using natural scene statistics and adaptive denoising
Anouar Kherchouche, Sid Ahmed Fezza, Wassim Hamidouche |
Neural Comput. Appl. | 2 |
| 2021 | HCiT: Deepfake Video Detection Using a Hybrid Model of CNN features and Vision TransformerabstractThe number of new falsified video contents is dramatically increasing, making the need to develop effective deepfake detection methods more urgent than ever. Even though many existing deepfake detection approaches show promising results, the majority of them still suffer from a number of critical limitations. In general, poor generalization results have been obtained under unseen or new deepfake generation methods. Consequently, in this paper, we propose a deepfake detection method called HCiT, which combines Convolutional Neural Network (CNN) with Vision Transformer (ViT). The HCiT hybrid architecture exploits the advantages of CNN to extract local information with the ViT's self-attention mechanism to improve the detection accuracy. In this hybrid architecture, the feature maps extracted from the CNN are feed into ViT model that determines whether a specific video is fake or real. Experiments were performed on Faceforensics++ and DeepFake Detection Challenge preview datasets, and the results show that the proposed method significantly outperforms the state-of-the-art methods. In addition, the HCiT method shows a great capacity for generalization on datasets covering various techniques of deepfake generation. The source code is available at: https://github.com/KADDAR-Bachir/HCiT Bachir Kaddar, Sid Ahmed Fezza, Wassim Hamidouche, Zahid Akhtar, Abdenour Hadid |
VCIP | 2 |
| 2021 | Light Field Image Coding Using VVC Standard and View Synthesis Based on Dual Discriminator GAN
Nader Bakir, Wassim Hamidouche, Sid Ahmed Fezza, Khouloud Samrouth, Olivier Déforges |
IEEE Trans. Multim. | 3 |
| 2020 | Light Field Image Coding Using Dual Discriminator Generative Adversarial Network And VVC Temporal ScalabilityabstractLight field technology represents a viable path for providing a high-quality VR content. However, such an imaging system generates a high amount of data leading to an urgent need for LF image compression solution. In this paper, we propose an efficient LF image coding scheme based on view synthesis. Instead of transmitting all the LF views, only some of them are coded and transmitted, while the remaining views are dropped. The transmitted views are coded using Versatile Video Coding (VVC) and used as reference views to synthesize the missing views at decoder side. The dropped views are generated using the efficient dual discriminator GAN model. The selection of reference/dropped views is performed using a rate distortion optimization based on the VVC temporal scalability. Experimental results show that the proposed method provides high coding performance and overcomes the state-of-the-art LF image compression solutions. Nader Bakir, Wassim Hamidouche, Sid Ahmed Fezza, Khouloud Samrouth, Olivier Déforges |
ICME | 3 |
| 2020 | Detection of Adversarial Examples in Deep Neural Networks with Natural Scene StatisticsabstractRecent studies have demonstrated that the deep neural networks (DNNs) are vulnerable to carefully-crafted perturbations added to a legitimate input image. Such perturbed images are called adversarial examples (AEs) and can cause DNNs to misclassify. Consequently, it is of paramount importance to develop detection methods of AEs, thus allowing to reject them. In this paper, we propose to characterize the AEs through the use of natural scene statistics (NSS). We demonstrate that these statistical properties are altered by the presence of adversarial perturbations. Based on this finding, we propose three different methods that exploit these scene statistics to determine if an input is adversarial or not. The proposed detection methods have been evaluated against four prominent adversarial attacks and on three standards datasets. The experimental results have shown that the proposed methods achieve a high detection accuracy while providing a low false positive rate. Anouar Kherchouche, Sid Ahmed Fezza, Wassim Hamidouche, Olivier Déforges |
IJCNN | 2 |
| 2020 | Extending 2D Saliency Models for Head Movement Prediction in 360-Degree Images using CNN-Based FusionabstractSaliency prediction can be of great benefit for 360-degree image/video applications, including compression, streaming, rendering and viewpoint guidance. It is therefore quite natural to adapt the 2D saliency prediction methods for 360-degree images. To achieve this, it is necessary to project the 360-degree image to 2D plane. However, the existing projection techniques introduce different distortions, which provides poor results and makes inefficient the direct application of 2D saliency prediction models to 360-degree content. Consequently, in this paper, we propose a new framework for effectively applying any 2D saliency prediction method to 360-degree images. The proposed framework particularly includes a novel convolutional neural network based fusion approach that provides more accurate saliency prediction while avoiding the introduction of distortions. The proposed framework has been evaluated with five 2D saliency prediction methods, and the experimental results showed the superiority of our approach compared to the use of weighted sum or pixel-wise maximum fusion methods. Ibrahim Djemai, Sid Ahmed Fezza, Wassim Hamidouche, Olivier Déforges |
ISCAS | 2 |
| 2020 | Natural Scene Statistics for Detecting Adversarial Examples in Deep Neural NetworksabstractThe deep neural networks (DNNs) have been adopted in a wide spectrum of applications. However, it has been demonstrated that their are vulnerable to adversarial examples (AEs): carefully-crafted perturbations added to a clean input image. These AEs fool the DNNs which classify them incorrectly. Therefore, it is imperative to develop a detection method of AEs allowing the defense of DNNs. In this paper, we propose to characterize the adversarial perturbations through the use of natural scene statistics. We demonstrate that these statistical properties are altered by the presence of adversarial perturbations. Based on this finding, we design a classifier that exploits these scene statistics to determine if an input is adversarial or not. The proposed method has been evaluated against four prominent adversarial attacks and on three standards datasets. The experimental results have shown that the proposed detection method achieves a high detection accuracy, even against strong attacks, while providing a low false positive rate. Anouar Kherchouche, Sid Ahmed Fezza, Wassim Hamidouche, Olivier Déforges |
MMSP | 2 |
| 2019 | RDO-Based Light Field Image Coding Using Convolutional Neural Networks and Linear ApproximationabstractThe increasing penetration of acquisition and display devices for Light Field (LF) content in the consumer market leads to the high proliferation of this new immersive media. This growing interest to LF images thus urgently raises the question of their compression. In this paper, we propose a convolutional neural networks (CNN)-based LF image coding scheme including both Rate Distortion Optimization (RDO) and post-processing steps. First, at the encoder side, the views are rearranged in sparse and dropped set of views. The former are compressed with a standard encoder and transmitted, while the dropped views are either linearly approximated or synthesized by a CNN using the encoded views as input. This choice is made on the basis of the proposed RDO process. At the decoder side, once the dropped views are either linearly approximated or synthesized by a CNN block, a post-processing step is performed to further enhance the quality of the reconstructed views. This post-processing block is based on superpixel to pixel-matching. Experimental results show that the proposed scheme provides views with high visual quality and overcomes the state-of-the-art LF image compression solutions by -30% in terms of BD-BR and 0.62 dB in BD-PSNR. Nader Bakir, Wassim Hamidouche, Olivier Déforges, Khouloud Samrouth, Sid Ahmed Fezza |
DCC | 5 |
| 2019 | Perceptual Evaluation of Adversarial Attacks for CNN-based Image ClassificationabstractDeep neural networks (DNNs) have recently achieved state-of-the-art performance and provide significant progress in many machine learning tasks, such as image classification, speech processing, natural language processing, etc. However, recent studies have shown that DNNs are vulnerable to adversarial attacks. For instance, in the image classification domain, adding small imperceptible perturbations to the input image is sufficient to fool the DNN and to cause misclassification. The perturbed image, called adversarial example, should be visually as close as possible to the original image. However, all the works proposed in the literature for generating adversarial examples have used the Lpnorms (L0, L2and L∞) as distance metrics to quantify the similarity between the original image and the adversarial example. Nonetheless, the Lpnorms do not correlate with human judgment, making them not suitable to reliably assess the perceptual similarity/fidelity of adversarial examples. In this paper, we present a database for visual fidelity assessment of adversarial examples. We describe the creation of the database and evaluate the performance of fifteen state-of-the-art full-reference (FR) image fidelity assessment metrics that could substitute Lpnorms. The database as well as subjective scores are publicly available to help designing new metrics for adversarial examples and to facilitate future research works. Sid Ahmed Fezza, Yassine Bakhti, Wassim Hamidouche, Olivier Déforges |
QoMEX | 1 |
| 2019 | Visual Security Assessment of Selective Video EncryptionabstractGiven the wide use of videos in various applications and across different devices, this raises the question of their security and confidentiality. In the last decade, many video encryption methods have been proposed in the literature. Accordingly, it becomes necessary to have a reliable assessment tool allowing evaluation of the efficiency of these video encryption methods, especially from the visual security point of view. Usually, the visual security is evaluated through the classical objective signal-based metrics. However, these metrics showed their limits as visual security metric, since they are not designed to deal with the security requirements, such as the determination of content intelligibility. Despite its obvious importance, very few visual security metrics have been proposed for the assessment of video encryption methods. This is mainly due to the lack of ground truth with subjective human scores for video encryption applications. In this paper, we present a new database for visual security assessment of selective video encryption. The database including unencrypted and encrypted video contents generated using different selective encryption schemes, as well as subjective scores, is publicly available to help designing new visual security metrics1. Sid Ahmed Fezza, Wassim Hamidouche, Reda Abdellah Kamraoui, Olivier Déforges |
QoMEX | 1 |
| 2017 | Using distortion and asymmetry determination for blind stereoscopic image quality assessment strategy
Sid Ahmed Fezza, Aladine Chetouani, Mohamed-Chaker Larabi |
J. Vis. Commun. Image Represent. | 1 |
| 2017 | Perceptually Driven Nonuniform Asymmetric Coding of Stereoscopic 3D VideoabstractAsymmetric stereoscopic video coding has already proven its effectiveness in reducing the bandwidth required for stereoscopic 3D delivery without degrading the visual quality. This approach, in which the left and right views are encoded with different levels of quality, relies on the perceptual theory of binocular suppression. However, to ensure comfortable 3D viewing, the just-noticeable level of asymmetry, i.e., the maximum quality gap between views, has to be carefully defined. Both subjectively and empirically fixed thresholds of asymmetry demonstrated either the maladjustment to content or dependency to the experimental design. This paper describes a new nonuniform asymmetric stereoscopic video coding method adaptively adjusting the level of asymmetry for each region of the image based on its perceptual significance. The proposed method uses a fully automated model that dynamically determines the best bounds of asymmetry for which the 3D viewing experience will not be altered. This is achieved by exploiting several human-visual-system-inspired models, namely, the binocular just-noticeable difference, and the visual saliency map and depth information. The simulation results show that the proposed method results in bit rate saving of up to 26% and provides better 3D visual quality compared with state-of-the-art asymmetric coding methods. Sid Ahmed Fezza, Mohamed-Chaker Larabi |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Asymmetric coding using Binocular Just Noticeable Difference and depth information for stereoscopic 3DabstractThe problem of determining the best level of asymmetry has been addressed by several recent works with the aim to guarantee an optimal binocular perception while keeping the minimum required information. To do so, subjective experiments have been conducted for the definition of an appropriate threshold. However, such an approach is lacking in terms of generalization because of the content variability. Moreover, using a fixed threshold does not allow an adaptation to the content and to the images' quality. The traditional asymmetric stereoscopic coding methods apply a uniform asymmetry by considering that all regions of an image have the same perceptual relevance which is not in compliance with the characteristics of human visual system (HVS). Consequently, this paper describes a fully automated model that dynamically determines the best bounds of asymmetry for each region of the image. Based on the Binocular Just Noticeable Difference (BJND) and the depth level in the scene, the proposed method achieves non-uniform reduction of spatial resolution of one view of the stereo pair with the aim to reduce bandwidth requirement. Experimental results show that the proposed method results in up to 43% of bitrate saving while outperforming the widely used asymmetric coding approaches in terms of 3D visual quality. Sid Ahmed Fezza, Mohamed-Chaker Larabi, Kamel Mohamed Faraoun |
ICASSP | 1 |
| 2014 | Stereoscopic image quality metric based on local entropy and binocular just noticeable differenceabstractDeveloping a metric that can reliably predict the perceptual 3D quality as perceived by the end user, is a challenging issue and a necessary tool for the success of 3D multimedia applications. The various attempts at predicting 3D quality of experience as the combination of 2D quality of the left and right images have shown their limitations, and particularly for the case of asymmetric distortions. In this paper we propose a full reference quality assessment metric for stereoscopic images based on the perceptual binocular characteristics. The proposed metric handles effectively the asymmetric distortions of stereoscopic images, by incorporating human visual system (HVS) characteristics. Our approach was motivated by the fact that in case of asymmetric quality, 3D perception mechanisms supports the view providing the most important and contrasted information. To achieve that, weighting factors are defined for the quality of each view according to the local information content. Add to that, to take into account the sensitivity of the HVS, quality score of each region are modulated based on the Binocular Just Noticeable Difference (BJND). Experimental results show that the proposed metric correlates better with human perception than the state-of-the-art metrics. Sid Ahmed Fezza, Mohamed-Chaker Larabi, Kamel Mohamed Faraoun |
ICIP | 1 |
| 2014 | Asymmetric coding of stereoscopic 3D based on perceptual significanceabstractAsymmetric stereoscopic coding is a very promising technique to decrease the bandwidth required for stereoscopic 3D delivery. However, one large obstacle is linked to the limit of asymmetric coding or the just noticeable threshold of asymmetry, so that 3D viewing experience is not altered. By way of subjective experiments, recent works have attempted to identify this asymmetry threshold. However, fixed threshold, highly dependent on the experiment design, do not allow to adapt to quality and content variation of the image. In this paper, we propose a new non-uniform asymmetric stereoscopic coding adjusting in a dynamic manner the level of asymmetry for each image region to ensure unaltered binocular perception. This is achieved by exploiting several HVS-inspired models; specifically we used the Binocular Just Noticeable Difference (BJND) combined with visual saliency map and depth information to quantify precisely the asymmetry threshold. Simulation results show that the proposed method results in up to 44% of bitrate saving and provides better 3D visual quality compared to state-of-the-art asymmetric coding methods. Sid Ahmed Fezza, Mohamed-Chaker Larabi, Kamel Mohamed Faraoun |
ICIP | 1 |
| 2014 | Feature-Based Color Correction of Multiview Video for Coding and Rendering EnhancementabstractMultiview video (MVV) consists of capturing the same scene with multiple cameras from different viewpoints. Therefore, substantial illumination and color inconsistencies can be observed between different views. These color mismatches can significantly reduce compression efficiency and rendering quality. In this paper, we propose a preprocessing method for correcting these color discrepancies in MVV. To consider the occlusion problem, our method is based on an improvement of histogram matching (HM) algorithm using only common regions across views. These regions are defined by an invariant feature detector (scale invariant feature transform), followed by random sample consensus algorithm to increase the matching robustness. In addition, to maintain temporal correlation, HM algorithm is applied on a temporal sliding window, allowing to cope with time-varying acquiring system, camera moving capture, and real-time broadcasting. Moreover, unlike always choosing the center view as the reference by default, we propose an automatic selection algorithm based on both views statistics and quality. The experimental results show that the proposed method increases coding efficiency with gains of up to 1.1 and 2.2 dB for the luminance and chrominance components, respectively. Furthermore, once the correction is performed, the color of real and rendered views is harmonized and looks very consistent as a whole. Sid Ahmed Fezza, Mohamed-Chaker Larabi, Kamel Mohamed Faraoun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |