EDBT 2026 Demo / reviewers in the wild / expert
Wassim Hamidouche
dblp:85/8759
· DBLP profile ↗
124ranked-venue papers
11as first author
56since 2021 · last 2026
0000-0002-0143-1756ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 105 · 10 first-author · 44 since 2021Artificial intelligence and machine learning · 11 · 10 since 2021Computer networks · 8 · 8 since 2021Systems, architecture and hardware · 6 · 1 since 2021Databases, data management, data science and information retrieval · 5Human-computer interaction and ubiquitous computing · 4Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African LanguagesabstractHao Yu, Tianyi Xu, Michael A. Hedderich, Wassim Hamidouche, Syed Waqas Zamir, David Ifeoluwa Adelani. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Michael A. Hedderich, Wassim Hamidouche, Syed Waqas Zamir, David Ifeoluwa Adelani |
ACL (1) | 4 |
| 2026 | Q-Codec: Reinforcement Learning based Adaptive Encoding for Real-Time Streaming of Immersive Aerial Videos
Mohit K. Sharma, Ibrahim Farhat, Wassim Hamidouche |
WCNC | 3 |
| 2026 | Complexity prediction of hardware and software video transcoding in the cloud
Taieb Chachou, Sid Ahmed Fezza, Wassim Hamidouche, Ghalem Belalem, Hadi Amirpour |
Multim. Tools Appl. | 3 |
| 2026 | Does data augmentation help or hinder the generalization of deepfake video detection?
Bachir Kaddar, Sid Ahmed Fezza, Elhocine Boutellaa, Wassim Hamidouche, Abdenour Hadid |
Multim. Tools Appl. | 4 |
| 2026 | Adversarial threats to vision transformers: evaluating robustness beyond CNNs
Ahmed Aldahdooh, Wassim Hamidouche, Olivier Déforges |
Neural Comput. Appl. | 2 |
| 2026 | Advancing Radio Map Construction and Obstacle Sensing: An Integrated Generative Framework in THz Band
Shuai Wang 0033, Yunhang Xie, Lingxiang Li, Zhi Chen 0002, Boyu Ning, Wassim Hamidouche, Lina Bariah, Samson Lasaulce, Mérouane Debbah |
IEEE Trans. Commun. | 7 |
| 2025 | DeeCLIP: A Robust and Generalizable Transformer-Based Framework for Detecting AI-Generated Images
Mamadou Keita, Wassim Hamidouche, Hessen Bougueffa Eutamene, Abdelmalik Taleb-Ahmed, Abdenour Hadid |
ACIVS | 2 |
| 2025 | Energy Backdoor Attack to Deep Neural NetworksabstractThe rise of deep learning (DL) has increased computing complexity and energy use, prompting the adoption of application specific integrated circuits (ASICs) for energy-efficient edge and mobile deployment. However, recent studies have demonstrated the vulnerability of these accelerators to energy attacks. Despite the development of various inference time energy attacks in prior research, backdoor energy attacks remain unexplored. In this paper, we design an innovative energy backdoor attack against deep neural networks (DNNs) operating on sparsity-based accelerators. Our attack is carried out in two distinct phases: backdoor injection and backdoor stealthiness. Experimental results using ResNet-18 and MobileNet-V2 models trained on CIFAR-10 and Tiny ImageNet datasets show the effectiveness of our proposed attack in increasing energy consumption on trigger samples while preserving the model’s performance for clean/regular inputs. This demonstrates the vulnerability of DNNs to energy backdoor attacks. The source code of our attack is available at: https://github.com/hbrachemi/energybackdoor. Hanene Brachemi Meftah, Wassim Hamidouche, Sid Ahmed Fezza, Olivier Déforges, Kassem Kallas |
ICASSP | 2 |
| 2025 | Can LLMs Revolutionize the Design of Explainable and Efficient TinyML Models?abstractThis paper introduces a novel framework for designing efficient neural network architectures specifically tailored to tiny machine learning (TinyML) platforms. By leveraging large language models (LLMs) for neural architecture search (NAS), a vision transformer (ViT)-based knowledge distillation (KD) strategy, and an explainability module, the approach strikes an optimal balance between accuracy, computational efficiency, and memory usage. The LLM-guided search explores a hierarchical search space, refining candidate architectures through Pareto optimization based on accuracy, multiply-accumulate operations (MACs), and memory metrics. The best-performing architectures are further fine-tuned using logits-based KD with a pre-trained ViT-B/16 model, which enhances generalization without increasing model size. Evaluated on the CIFAR-100 dataset and deployed on an STM32H7 microcontroller (MCU), the three proposed models, LMaNet-Elite, LMaNet-Core, and QwNet-Core, achieve accuracy scores of 74.50%, 74.20% and 73.00%, respectively. All three models surpass current state-of-the-art (SOTA) models, such as MCUNet-in3/in4 (69.62% / 72.86%) and XiNet (72.27%), while maintaining a low computational cost of less than 100 million MACs and adhering to the stringent 320 KB static random-access memory (SRAM) constraint. These results demonstrate the efficiency and performance of the proposed framework for TinyML platforms, underscoring the potential of combining LLM-driven search, Pareto optimization, KD, and explainability to develop accurate, efficient, and interpretable models. This approach opens new possibilities in NAS, enabling the design of efficient architectures specifically suited for TinyML. To facilitate further research and development in this field, the proposed framework and the best-performing architectures are made publicly available at Link. Christophe El Zeinaty, Wassim Hamidouche, Glenn Herrou, Daniel Ménard, Mérouane Debbah |
IJCNN | 2 |
| 2025 | Bi-LORA: A Vision-Language Approach for Synthetic Image DetectionabstractABSTRACT Advancements in deep image synthesis techniques, such as generative adversarial networks (GANs) and diffusion models (DMs), have ushered in an era of generating highly realistic images. While this technological progress has captured significant interest, it has also raised concerns about the high challenge in distinguishing real images from their synthetic counterparts. This paper takes inspiration from the potent convergence capabilities between vision and language, coupled with the zero‐shot nature of vision‐language models (VLMs). We introduce an innovative method called Bi‐LORA that leverages VLMs, combined with low‐rank adaptation (LORA) tuning techniques, to enhance the precision of synthetic image detection for unseen model‐generated images. The pivotal conceptual shift in our methodology revolves around reframing binary classification as an image captioning task, leveraging the distinctive capabilities of cutting‐edge VLM, notably bootstrapping language image pre‐training (BLIP)2. Rigorous and comprehensive experiments are conducted to validate the effectiveness of our proposed approach, particularly in detecting unseen diffusion‐generated images from unknown diffusion‐based generative models during training, showcasing robustness to noise, and demonstrating generalisation capabilities to GANs. The experiments show that Bi‐LORA outperforms state of the art models in cross‐generator tasks because it leverages multi‐modal learning, open‐world visual knowledge, and benefits from robust, high‐level semantic understanding. By combining visual and textual knowledge, it can handle variations in the data distribution (such as those caused by different generators) and maintain strong performance across different domains. Its ability to transfer knowledge, robustly extract features and perform zero‐shot learning also contributes to its generalisation capabilities, making it more adaptable to new generators. The experimental results showcase an impressive average accuracy of 93.41% in synthetic image detection on unseen generation models. The code and models associated with this research can be publicly accessed at https://github.com/Mamadou‐Keita/VLM‐DETECT . Mamadou Keita, Wassim Hamidouche, Hessen Bougueffa Eutamene, Abdelmalik Taleb-Ahmed, David Camacho, Abdenour Hadid |
Expert Syst. J. Knowl. Eng. | 2 |
| 2025 | 360-degree video super resolution and quality enhancement challenge: Methods and results
Ahmed Telili, Wassim Hamidouche, Ibrahim Farhat, Hadi Amirpour, Christian Timmerer, Ibrahim Khadraoui, Jiajie Lu, The Van Le, Jeonneung Baek, Yiying Wei, Jiancheng Huang |
Signal Process. Image Commun. | 2 |
| 2025 | Convex Hull Prediction Methods for Bitrate Ladder Construction: Design, Evaluation, and ComparisonabstractHTTP adaptive streaming (HAS) has emerged as a prevalent approach for over-the-top (OTT) video streaming services due to its ability to deliver a seamless user experience. A fundamental component of HAS is the bitrate ladder, which comprises a set of encoding parameters (e.g., bitrate-resolution pairs) used to encode the source video into multiple representations. This adaptive bitrate ladder enables the client’s video player to dynamically adjust the quality of the video stream in real-time based on fluctuations in network conditions, ensuring uninterrupted playback by selecting the most suitable representation for the available bandwidth. The most straightforward approach involves using a fixed bitrate ladder for all videos, consisting of pre-determined bitrate-resolution pairs known as one-size-fits-all . Conversely, the most reliable technique relies on intensively encoding all resolutions over a wide range of bitrates to build the convex hull , thereby optimizing the bitrate ladder by selecting the representations from the convex hull for each specific video. Several techniques have been proposed to predict content-based ladders without performing a costly, exhaustive search encoding. This article provides a comprehensive review of various convex hull prediction methods, including both conventional and learning-based approaches. Furthermore, we conduct a benchmark study of several handcrafted- and deep learning (DL)-based approaches for predicting content-optimized convex hulls across multiple codec settings. The considered methods are evaluated on our proposed large-scale dataset, which includes 300 UHD video shots encoded with software and hardware encoders using three state-of-the-art video standards, including AVC/H.264, HEVC/H.265, and VVC/H.266, at various bitrate points. Our analysis provides valuable insights and establishes baseline performance for future research in this field ( Dataset URL : https://nasext-vaader.insa-rennes.fr/ietr-vaader/datasets/br_ladder ). Ahmed Telili, Wassim Hamidouche, Hadi Amirpour, Sid Ahmed Fezza, Christian Timmerer, Luce Morin |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Dicetrack: Lightweight Dice Classification on Resource-Constrained Platforms with Optimized Deep Learning ModelsabstractThis paper introduces DiceTrack, an innovative Deep Learning (DL) platform for detecting dice in board games. Deploying robust models on microcontrollers (MCUs) presents challenges due to memory and computational constraints. We focus on optimizing MobileNet for seamless ESP32 deployment and propose two novel ultra-lightweight models, Separable Convolutional Layers with Quantization Network (SCLQNet) and Binarized Neural Network (BNNet), for dice classification. SCLQNet uses separable convolutional layers with quantized weights and activations, while BNNet employs a unique Binarized architecture. Further, we create DiceVision, a custom dice classification dataset tailored for real-time digital board games. Comprehensive evaluations on ESP32 and Raspberry Pi 4 showcase the efficiency of the proposed models. SCLQNet and BNNet achieve 97.5% and 97.4% accuracy with 32KB and 22KB model sizes. Notably, SCLQNet takes only 402 ms for ESP32 inference, enabling 88% latency reduction compared to quantized MobileNet. Christophe El Zeinaty, Glenn Herrou, Wassim Hamidouche, Daniel Ménard |
ICASSP | 3 |
| 2024 | Res-NeRV: Residual Blocks For A Practical Implicit Neural Video DecoderabstractThis paper proposes the integration of residual blocks into neural representation for videos (NeRV)-based architectures with the aim of enhancing the reconstruction of detailed patterns and high-level features. Additionally, a coding pipeline is introduced, placing the implicit neural decoder in a real-life video streaming framework. Indeed, DeepCABAC is employed for model compression, applying a quantization scheme followed by the context-adaptive binary arithmetic coding (CABAC) entropy coding algorithm, ultimately leading to bitstream generation. Our method outperforms NeRV, as well as x264 and x265, achieving BD-rate gains against NeRV: $-12.06 \%$ using PSNR and $-14.25 \%$ using MS-SSIM. Furthermore, it exhibits superior subjective quality compared to NeRV, attributed to enhanced high-level feature reconstruction. This observed behavior encourages the application of our method to other NeRV-based models, such as E-NeRV. Marwa Tarchouli, Thomas Guionnet, Marc Rivière, Wassim Hamidouche, Meriem Outtas, Olivier Déforges |
ICIP | 4 |
| 2024 | ODVISTA: An Omnidirectional Video Dataset for Super-Resolution and Quality Enhancement TasksabstractOmnidirectional or 360-degree video is being increasingly deployed, largely due to the latest advancements in immersive virtual reality (VR) and extended reality (XR) technology. However, the adoption of these videos in streaming encounters challenges related to bandwidth and latency, particularly in mobility conditions such as with unmanned aerial vehicles (UAVs). Adaptive resolution and compression aim to preserve quality while maintaining low latency under these constraints, yet downscaling and encoding can still degrade quality and introduce artifacts. Machine learning (ML)-based super-resolution (SR) and quality enhancement techniques offer a promising solution by enhancing detail recovery and reducing compression artifacts. However, current publicly available 360-degree video SR datasets lack compression artifacts, which limit research in this field. To bridge this gap, this paper introduces omnidirectional video streaming dataset (ODVista), which comprises 200 high-resolution and high-quality videos downscaled and encoded at four bitrate ranges using the high-efficiency video coding (HEVC)/H.265 standard. Evaluations show that the dataset not only features a wide variety of scenes but also spans different levels of content complexity, which is crucial for robust solutions that perform well in real-world scenarios and generalize across diverse visual environments. Additionally, we evaluate the performance, considering both quality enhancement and runtime, of two handcrafted and two ML-based SR models on the validation and testing sets of ODVista. Dataset URL: https://github.com/Omnidirectional-video-group/ODVista Ahmed Telili, Ibrahim Farhat, Wassim Hamidouche, Hadi Amirpour |
ICIP | 3 |
| 2024 | FIDAVL: Fake Image Detection and Attribution Using Vision-Language Model
Mamadou Keita, Wassim Hamidouche, Hessen Bougueffa Eutamene, Abdelmalik Taleb-Ahmed, Abdenour Hadid |
ICPR (21) | 2 |
| 2024 | EVCA: Enhanced Video Complexity AnalyzerabstractThe optimization of video compression and streaming workflows critically relies on understanding the video complexity, including both spatial and temporal features. These features play a vital role in guiding rate control, predicting video encoding parameters (such as resolution and frame rate), and selecting test videos for subjective analysis. Traditional methods primarily utilize Spatial Information (SI) and Temporal Information (TI) to measure these spatial and temporal complexity features, respectively. Moreover, the Video Complexity Analyzer (VCA) has been introduced as a tool employing Discrete Cosine Transform (DCT)-based functions, namely E and h, to evaluate the spatial and temporal complexity features, respectively. In this paper, we introduce the Enhanced Video Complexity Analyzer (EVCA), an advanced tool that integrates the functionalities of both VCA and the SITI approach. Developed in Python to ensure compatibility with GPU processing, EVCA enhances the definition of temporal complexity originally used in VCA. This refinement significantly improves the detection of temporal complexity features in VCA (i.e., h), raising its Pearson Correlation Coefficient (PCC) from 0.6 to 0.77. Furthermore, EVCA demonstrates exceptional performance on Graphics Processing Unit (GPU) devices, achieving feature extraction speeds exceeding 1200 fps for 1080p resolution videos. Hadi Amirpour, Mohammad Ghasempour, Lingfeng Qu, Wassim Hamidouche, Christian Timmerer |
MMSys | 4 |
| 2024 | MVCD: Multi-Dimensional Video Compression DatasetabstractIn the field of video streaming, the optimization of video encoding and decoding processes is crucial for delivering high-quality video content. Given the growing concern about carbon dioxide emissions, it is equally necessary to consider the energy consumption associated with video streaming. Therefore, to take advantage of machine learning techniques for optimizing video delivery, a dataset encompassing the energy consumption of the encoding and decoding process is needed. This paper introduces a comprehensive dataset featuring diverse video content, encoded and decoded using various codecs and spanning different devices. The dataset includes 1000 videos encoded with four resolutions (2160p, 1080p, 720p, and 540p) at two frame rates (30fps and 60fps), resulting in eight unique encodings for each video. Each video is further encoded with four different codecs — AVC (libx264), HEVC (libx265), AV1 (libsvtav1), and VVC (VVenC) — at four quality levels defined by QPs of 22, 27, 32 and 37. In addition, for AV1, three additional QPs of 35, 46 and 55 are considered. We measure both encoding and decoding time and energy consumption on various devices to provide a comprehensive evaluation, employing various metrics and tools. Additionally, we assess encoding bitrate and quality using quality metrics such as PSNR, SSIM, MS-SSIM, and VMAF. All data and the reproduction commands and scripts have been made publicly available as part of the dataset, which can be used for various applications such as rate and quality control, resource allocation, and energy-efficient streaming.Dataset URL: https://github.com/cd-athena/MVCD. Hadi Amirpour, Mohammad Ghasempour, Farzad Tashtarian, Ahmed Telili, Samira Afzal, Wassim Hamidouche, Christian Timmerer |
VCIP | 6 |
| 2024 | NeRV++: An Enhanced Implicit Neural Video RepresentationabstractNeural fields, also known as implicit neural representations (INRs), have shown a remarkable capability of representing, generating, and manipulating various data types, allowing for continuous data reconstruction at a low memory footprint. Though promising, INRs applied to video compression still need to improve their rate-distortion performance by a large margin, and require a huge number of parameters and long training iterations to capture high-frequency details, limiting their wider applicability. Resolving this problem remains a quite challenging task, which would make INRs more accessible in compression tasks. We take a step towards resolving these shortcomings by introducing neural representations for videos (NeRV)++, an enhanced implicit neural video representation (INVR), as more straightforward yet effective enhancement over the original NeRV decoder architecture, featuring separable conv2d residual blocks (SCRBs) that sandwiches the upsampling block (UB), and a bilinear interpolation skip layer for improved feature representation. NeRV++ allows videos to be directly represented as a function approximated by a neural network, and significantly enhance the representation capacity beyond current INR-based video codecs. We evaluate our method on UVG, MCL JVC, and Bunny datasets, achieving competitive results for video compression with INRs. This achievement narrows the gap to autoencoder-based video coding, marking a significant stride in INR-based video compression research. The code of NeRV++ is available on GitHub. Ahmed Ghorbel, Wassim Hamidouche, Luce Morin |
VCIP | 2 |
| 2024 | Error Concealment Capacity Analysis of Standard Video Encoders for Real-Time Immersive Video Streaming from UAVsabstractIn this work, we analyze the error concealment capacity of various video codecs, for real-time streaming of 360° videos from a unmanned aerial vehicle (UAV) to ground based user. For this study, we evaluate the performance of various recent encoders in ISO/IEC and ITU-T standards: high-efficiency video coding (HEVC)/H.265 and advanced video coding (AVC)/H.264, including both their software and hardware implementations. We also consider the AOMedia video 1 (AV1) encoder from alliance for open media (AOM) formats. We benchmark the codecs by examining the video freezing time for each codec, i.e., the number of dropped frames, for transmission of 360° videos over an air-to-ground (A2G) wireless channel without any forward error correction. Interestingly, our results show that while AV1 achieves the best quality in terms of peak signal-to-noise ratio (PSNR), it exhibits worst error concealment capacity, resulting in the longest video freezing time among all the encoders. Further, we observe that the AVC yields the most robust performance against the impairments induced by the wireless channel. To present a complete picture, we also benchmark the codecs with respect to coding efficiency, and encoding/decoding latency at multiple bit-rates. In this study, we find that the AVC yields lowest latency for both software and hardware implementations. However, considering video quality, the hardware-based implementation of HEVC provides the best trade-off among all the parameters. Overall, our results provide novel and interesting insights on the choice of encoders for real-time 360° video transmission from a UAV to ground user. Mohit K. Sharma, Ibrahim Farhat, Wassim Hamidouche |
WCNC | 3 |
| 2024 | Strategic safeguarding: A game theoretic approach for analyzing attacker-defender behavior in DNN backdoorsabstractDeep neural networks (DNNs) are fundamental to modern applications like face recognition and autonomous driving. However, their security is a significant concern due to various integrity risks, such as backdoor attacks. In these attacks, compromised training data introduce malicious behaviors into the DNN, which can be exploited during inference or deployment. This paper presents a novel game-theoretic approach to model the interactions between an attacker and a defender in the context of a DNN backdoor attack. The contribution of this approach is multifaceted. First, it models the interaction between the attacker and the defender using a game-theoretic framework. Second, it designs a utility function that captures the objectives of both parties, integrating clean data accuracy and attack success rate. Third, it reduces the game model to a two-player zero-sum game, allowing for the identification of Nash equilibrium points through linear programming and a thorough analysis of equilibrium strategies. Additionally, the framework provides varying levels of flexibility regarding the control afforded to each player, thereby representing a range of real-world scenarios. Through extensive numerical simulations, the paper demonstrates the validity of the proposed framework and identifies insightful equilibrium points that guide both players in following their optimal strategies under different assumptions. The results indicate that fully using attack or defense capabilities is not always the optimal strategy for either party. Instead, attackers must balance inducing errors and minimizing the information conveyed to the defender, while defenders should focus on minimizing attack risks while preserving benign sample performance. These findings underscore the effectiveness and versatility of the proposed approach, showcasing optimal strategies across different game scenarios and highlighting its potential to enhance DNN security against backdoor attacks. Kassem Kallas, Quentin Le Roux, Wassim Hamidouche, Teddy Furon |
EURASIP J. Inf. Secur. | 3 |
| 2024 | Hierarchical Learning and Dummy Triplet Loss for Efficient Deepfake DetectionabstractThe advancement of generative models has made it easier to create highly realistic Deepfake videos. This accessibility has led to a surge in research on Deepfake detection to mitigate potential misuse. Typically, Deepfake detection models utilize binary backbones, even though the training dataset contains additional exploitable information, such as the Deepfake generation method employed for each video. However, recent findings suggest that inferring a binary class from a multi-class backbone yields superior performance compared to directly employing a binary backbone. Building upon this research, our article introduces two novel methods to infer a binary class from a multi-class backbone. The first method, named root dummies , leverages the dummy triplet loss, which employs fixed vectors (i.e., dummies) instead of mined positives and negatives in the triplet loss. By training the multi-class backbone with these dummies, we can easily infer a binary class during testing by adjusting the number of dummies (from six during training to two during inference). Through this approach, we achieve an accuracy improvement of 0.23% compared to the existing inference method, without requiring additional training. The second proposed method is transfer learning. It involves training a classifier, such as a support vector machine, to predict binary classes based on the image embeddings generated by the multi-class backbone. Although this method necessitates additional training, it further enhances the model’s performance, resulting in an accuracy increase of 1.79%. In summary, our proposed methods improve the accuracy of Deepfake detection by simply modifying the number of classes during training, making them suitable for integration into a variety of existing Deepfake training pipelines. Additionally, to foster reproducible research, we have made the source code of our solution publicly available at https://github.com/beuve/DmyT . Nicolas Beuve, Wassim Hamidouche, Olivier Déforges |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Deepfake Detection Using Spatiotemporal TransformerabstractRecent advances in generative models and the availability of large-scale benchmarks have made deepfake video generation and manipulation easier. Nowadays, the number of new hyper-realistic deepfake videos used for negative purposes is dramatically increasing, thus creating the need for effective deepfake detection methods. Although many existing deepfake detection approaches, particularly CNN-based methods, show promising results, they suffer from several drawbacks. In general, poor generalization results have been obtained under unseen/new deepfake generation methods. The crucial reason for the above defect is that CNN-based methods focus on the local spatial artifacts, which are unique for every manipulation method. Therefore, it is hard to learn the general forgery traces of different manipulation methods without considering the dependencies that extend beyond the local receptive field. To address this problem, this article proposes a framework that combines Convolutional Neural Network (CNN) with Vision Transformer (ViT) to improve detection accuracy and enhance generalizability. Our method, namedHCiT, exploits the advantages of CNNs to extract meaningful local features, as well as the ViT’s self-attention mechanism to learn discriminative global contextual dependencies in a frame-level image explicitly. In this hybrid architecture, the high-level feature maps extracted from the CNN are fed into the ViT model that determines whether a specific video is fake or real. Experiments were performed on Faceforensics++, DeepFake Detection Challenge preview, Celeb datasets, and the results show that the proposed method significantly outperforms the state-of-the-art methods. In addition, the HCiT method shows a great capacity for generalization on datasets covering various techniques of deepfake generation. The source code is available at: https://github.com/KADDAR-Bachir/HCiT Bachir Kaddar, Sid Ahmed Fezza, Zahid Akhtar, Wassim Hamidouche, Abdenour Hadid, Joan Serra-Sagristà |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | 2BiVQA: Double Bi-LSTM-based Video Quality Assessment of UGC VideosabstractRecently, with the growing popularity of mobile devices as well as video sharing platforms (e.g., YouTube, Facebook, TikTok, and Twitch), User-Generated Content (UGC) videos have become increasingly common and now account for a large portion of multimedia traffic on the internet. Unlike professionally generated videos produced by filmmakers and videographers, typically, UGC videos contain multiple authentic distortions, generally introduced during capture and processing by naive users. Quality prediction of UGC videos is of paramount importance to optimize and monitor their processing in hosting platforms, such as their coding, transcoding, and streaming. However, blind quality prediction of UGC is quite challenging, because the degradations of UGC videos are unknown and very diverse, in addition to the unavailability of pristine reference. Therefore, in this article, we propose an accurate and efficient Blind Video Quality Assessment (BVQA) model for UGC videos, which we name 2BiVQA for double Bi-LSTM Video Quality Assessment. 2BiVQA metric consists of three main blocks, including a pre-trained Convolutional Neural Network to extract discriminative features from image patches, which are then fed into two Recurrent Neural Networks for spatial and temporal pooling. Specifically, we use two Bi-directional Long Short-term Memory networks, the first is used to capture short-range dependencies between image patches, while the second allows capturing long-range dependencies between frames to account for the temporal memory effect. Experimental results on recent large-scale UGC VQA datasets show that 2BiVQA achieves high performance at lower computational cost than most state-of-the-art VQA models. The source code of our 2BiVQA metric is made publicly available at https://github.com/atelili/2BiVQA . Ahmed Telili, Sid Ahmed Fezza, Wassim Hamidouche, Hanene Brachemi Meftah |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | AICT: An Adaptive Image Compression TransformerabstractMotivated by the efficiency investigation of the Tranformer-based transform coding framework, namely SwinT-ChARM, we propose to enhance the latter, as first, with a more straightforward yet effective Tranformer-based channel-wise auto-regressive prior model, resulting in an absolute image compression transformer (ICT). Current methods that still rely on ConvNet-based entropy coding are limited in long-range modeling dependencies due to their local connectivity and an increasing number of architectural biases and priors. On the contrary, the proposed ICT can capture both global and local contexts from the latent representations and better parameterize the distribution of the quantized latents. Further, we leverage a learnable scaling module with a sandwich ConvNeXt-based pre/post-processor to accurately extract more compact latent representation while reconstructing higher-quality images. Extensive experimental results on benchmark datasets showed that the proposed adaptive image compression transformer (AICT) framework significantly improves the trade-off between coding efficiency and decoder complexity over the versatile video coding (VVC) reference encoder (VTM-18.0) and the neural codec SwinT-ChARM. Ahmed Ghorbel, Wassim Hamidouche, Luce Morin |
ICIP | 2 |
| 2023 | Efficient Per-Shot Transformer-Based Bitrate Ladder Prediction for Adaptive Video StreamingabstractRecently, HTTP adaptive streaming (HAS) has become a standard approach for over-the-top (OTT)-based video streaming services due to its ability to provide smooth streaming. In HAS, stream representations are encoded to target a specific bitrate providing a wide range of operating bitrates known as the bitrate ladder. In the past, a fixed bitrate ladder approach for all videos has been widely used. However, such a method does not consider video content, which can vary considerably in motion, texture, and scene complexity. Moreover, building a per-title bitrate ladder based on an exhaustive encoding is quite expensive due to the large encoding parameter space. Thus, alternative solutions allowing accurate and efficient per-title bitrate ladder prediction are in great demand. On the other hand, self-attention-based architectures have achieved tremendous performance in large language models (LLMs) and particularly vision transformers (ViTs) in computer vision tasks. Therefore, this paper investigates ViT’s capabilities in building an efficient bitrate ladder without performing any encoding process. We provide the first in-depth analysis of the prediction accuracy and the complexity overhead induced by the ViTs model in predicting the bitrate ladder on a large and diverse video dataset. The source code of the proposed solution and the dataset will be made publicly available. Ahmed Telili, Wassim Hamidouche, Sid Ahmed Fezza, Luce Morin |
ICIP | 2 |
| 2023 | Live and low energy VVC Video Decoding powered by the OpenVVC Decoder on ARM PlatformabstractThis demonstration showcases the potential of open-source software implementation for the new versatile video coding (VVC) standard, OpenVVC. The most complex VVC tools were optimized for ARM-type architectures using data parallelism through SIMD instructions. The demonstration has been tested on the NVIDIA Jetson AGX Xavier and on a NVIDIA SHIELD Android TV which showcased real-time decoding for FHD and HD video resolutions. By combining extensive data level parallelism with frame level parallelism, OpenVVC is able to maintain a remarkably low memory and energy consumption while achieving real-time decoding of videos with high resolution. These features present a great advantage for its integration on embedded devices with low computing and memory resources. Ibrahim Farhat, Pierre-Loup Cabarat, Wassim Hamidouche, Patrice Angot, Philippe Gonon, Daniel Ménard |
ISCAS | 3 |
| 2023 | Selective Secret Sharing Scheme for Privacy of Image and Video Compressed in MPEG-Like FormatsabstractThis paper presents a reliable image and video secret storing method, aiming at ensuring the protected storage of multimedia files without risk for their owner to have their privacy violated. It relies on a particular property (standardized in MPEG-A Part.21 VIMAF) of video and image standards such as H.264/AVC, H.265/HEVC and MPEG-HEIF: a non negligible portion of the bitstream (so-called cipherable bits) can be altered without modifying the decodability of the stream. Such a modification results in an image or a video that is both a valid and visually encrypted media file. Not willing to be bothered by a complex key management system, we take inspiration from Shamir's secret sharing scheme, and propose a method where the secret to be shared is the original (unciphered) set of cipherable bits, that are encoded and divided in$n$shares, among which$k-1$or less do not permit to reconstruct the secret, while$k$or more allow to recover it. Those$n$shares are then used to generate$n$ciphered fully standard compliant versions of the media. The rightful owner of the file will easily store these$n$versions in various Cloud storage locations, and recover them all when the fully deciphered file will be needed, while a hacker will most generally obtain only some of those files, hence not be capable to decrypt the file. Cyril Bergeron, Catherine Lamy-Bergot, Wassim Hamidouche, William Puech |
MMSP | 3 |
| 2023 | Energy Consumption and Carbon Footprint of Modern Video Decoding SoftwareabstractThe estimation of energy consumption has become vital in developing eco-friendly and sustainable video streaming solutions to monitor CO2 emissions. In this paper, we seek to evaluate and compare the energy consumption and CO2 emissions of the decoding process related to three popular video coding standards, namely AVC, HEVC, VVC, along with two video formats VP9, and AV1 through their real-time software decoders, including h264, hevc, VVdeC/OpenVVC, vp9, and libdav1d. The evaluation is conducted on two types of consumer hardware, desktop PC and laptop. To ensure a fair evaluation, we also assess the coding efficiency of software encoder implementations using three objective quality metrics. The experimental results revealed that the h264 decoder consumes the lowest energy and is associated with the lowest CO2 emissions compared to other decoders on both hardware platforms. On the other hand, the VVenC encoder enhances coding efficiency at the cost of increased decoding energy consumption and CO2 emissions, particularly noticeable in the case of the OpenVVC decoder. Meanwhile, x265/hevc achieves a compelling balance between coding efficiency and decoding energy consumption. The full results of this work are available at https://decodingenergy.github.io/decoding_energy_co2.html. Taieb Chachou, Wassim Hamidouche, Sid Ahmed Fezza, Ghalem Belalem |
MMSP | 2 |
| 2023 | Energy Efficient VVC Decoding on Mobile PlatformabstractRecently, global demand for high-resolution videos and new multimedia applications have created the need for a new video coding standard. Hence, in July 2020 the Versatile Video Coding (VVC) standard was released providing up to 40% bit-rate saving for the same video quality compared to its predecessor High Efficiency Video Coding (HEVC). However, this bit-rate saving comes at the cost of high computational complexity, particularly for live applications and on resource-constraint embedded devices. This paper presents an power-efficient VVC decoder implementation designed for low-resource platforms. This latter exploits optimization techniques such as data level parallelism using Single Instruction Multiple Data (SIMD) instructions and functional level parallelism using frame, tile and slice-based parallelisms. The results showed that the OpenVVC decoder achieve real-time decoding of Full High Definition (FHD) resolution at 30 fps targeting a platform with 8 cores with a maximum frequency of 2.2 Ghz and High Definition (HD) real-time decoding at 30 fps for platforms using 4 cores with a maximum frequency of 1.8 Ghz. In terms of average consumed power, OpenVVC showed around 5.6 watts and 1.7 watts for the 8 and 4 cores platforms, respectively. In addition, it comes with the best trade-off between what is achievable in real-time and the power consumed in comparison to the state-of-the-art implementation. Ibrahim Farhat, Pierre-Loup Cabarat, Daniel Ménard, Wassim Hamidouche, Olivier Déforges |
MMSP | 4 |
| 2023 | Open-Source Toolkit for Live End-to-End 4K VVC Intra CodingabstractVersatile Video Coding (VVC/H.266) takes video coding to the next level by doubling the coding efficiency over its predecessors for the same subjective quality, but at the cost of immense coding complexity. Therefore, VVC calls for aggressively optimized codecs to make it feasible for live streaming media applications. This paper introduces the first public end-to-end (E2E) pipeline for live 4K30p VVC intra coding and streaming. The pipeline is made up of three open-source components: 1) uvg266 for VVC encoding; 2) uvgRTP for VVC streaming; and 3) OpenVVC for VVC decoding. The proposed setup is demonstrated with a proof-of-concept prototype that implements the encoder end on AMD ThreadRipper 2990WX and the decoder end on Nvidia Jetson AGX Orin. Our prototype is almost 34 000 times as fast as the corresponding E2E pipeline built around the VTM codec. Respectively, it achieves 3.3 times speedup without any significant coding overhead over the pipeline that utilizes the fastest possible configuration of the well-known VVenC/VVdeC codec. These results indicate that our prototype is currently the only viable open-source solution for live 4K VVC intra coding and streaming. Marko Viitanen, Joose Sainio, Alexandre Mercat, Guillaume Gautier, Jarno Vanne, Ibrahim Farhat, Pierre-Loup Cabarat, Wassim Hamidouche, Daniel Ménard |
MMSys | 8 |
| 2023 | Revisiting model's uncertainty and confidences for adversarial example detection
Ahmed Aldahdooh, Wassim Hamidouche, Olivier Déforges |
Appl. Intell. | 2 |
| 2023 | Complexity assessment of the intra prediction in Versatile Video Coding
Naima Zouidi, Amina Kessentini, Wassim Hamidouche, Nouri Masmoudi, Daniel Ménard |
Multim. Tools Appl. | 3 |
| 2023 | Machine Learning Based Efficient QT-MTT Partitioning Scheme for VVC Intra EncodersabstractThe next-generation Versatile Video Coding (VVC) standard introduces a new Multi-Type Tree (MTT) block partitioning structure that supports Binary-Tree (BT) and Ternary-Tree (TT) splits in both vertical and horizontal directions. This new approach leads to five possible splits at each block depth. It thereby improves the coding efficiency of VVC over that of the preceding High Efficiency Video Coding (HEVC) standard, which only supports Quad-Tree (QT) partitioning with a single split per block depth. However, MTT also has brought a considerable impact on encoder computational complexity. This paper proposes a two-stage learning-based technique to tackle the complexity overhead of MTT in VVC intra encoders. In our scheme, the input block is first processed by a Convolutional Neural Network (CNN) to predict its spatial features through a vector of probabilities describing the partition at each$4\times 4$edge. Subsequently, a Decision Tree (DT) model leverages this vector of spatial features to predict the most likely splits at each block. Finally, based on this prediction, only the$N$most likely splits are processed by the Rate-Distortion (RD) process of the encoder. In order to train our CNN and DT models on a wide range of image contents, we also propose a public VVC frame partitioning dataset based on existing image dataset encoded with the VVC reference software encoder. Our solution relying on the top-3 configuration reaches 47.4% complexity reduction for a negligible bitrate increase of 0.79%. A top-2 configuration enables a higher complexity reduction of 70.4% for 2.49% bitrate loss. These results emphasize a better trade-off between VTM intra-coding efficiency and complexity reduction compared to the state-of-the-art solutions. The source code of the proposed method and the training dataset are made publicly available at GitHub. Alexandre Tissier, Wassim Hamidouche, Souhaiel Belhadj Dit Mdalsi, Jarno Vanne, Franck Galpin, Daniel Ménard |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Deep-Based Film Grain Removal and SynthesisabstractIn this paper, deep learning-based techniques for film grain removal and synthesis that can be applied in video coding are proposed. Film grain is inherent in analog film content because of the physical process of capturing images and video on film. It can also be present in digital content where it is purposely added to reflect the era of analog film and to evoke certain emotions in the viewer or enhance the perceived quality. In the context of video coding, the random nature of film grain makes it both difficult to preserve and very expensive to compress. To better preserve it while compressing the content efficiently, film grain is removed and modeled before video encoding and then restored after video decoding. In this paper, a film grain removal model based on an encoder-decoder architecture and a film grain synthesis model based on a conditional generative adversarial network (cGAN) are proposed. Both models are trained on a large dataset of pairs of clean (grain-free) and grainy images. Quantitative and qualitative evaluations of the developed solutions were conducted and showed that the proposed film grain removal model is effective in filtering film grain at different intensity levels using two configurations: 1) a non-blind configuration where the film grain level of the grainy input is known and provided as input; and 2) a blind configuration where the film grain level is unknown. As for the film grain synthesis task, the experimental results show that the proposed model is able to reproduce realistic film grain with a controllable intensity level specified as input. Zoubida Ameur, Wassim Hamidouche, Edouard François, Milos Radosavljevic, Daniel Ménard, Claire-Hélène Demarty |
IEEE Trans. Image Process. | 2 |
| 2023 | Predictive Uncertainty Estimation for Camouflaged Object DetectionabstractUncertainty is inherent in machine learning methods, especially those for camouflaged object detection aiming to finely segment the objects concealed in background. The strong enquote center bias of the training dataset leads to models of poor generalization ability as the models learn to find camouflaged objects around image center, which we define as enquote model bias. Further, due to the similar appearance of camouflaged object and its surroundings, it is difficult to label the accurate scope of the camouflaged object, especially along object boundaries, which we term as enquote data bias. To effectively model the two types of biases, we resort to uncertainty estimation and introduce predictive uncertainty estimation technique, which is the sum of model uncertainty and data uncertainty, to estimate the two types of biases simultaneously. Specifically, we present a predictive uncertainty estimation network (PUENet) that consists of a Bayesian conditional variational auto-encoder (BCVAE) to achieve predictive uncertainty estimation, and a predictive uncertainty approximation (PUA) module to avoid the expensive sampling process at test-time. Experimental results show that our PUENet achieves both highly accurate prediction, and reliable uncertainty estimation representing the biases within both model parameters and the datasets. Yi Zhang 0076, Jing Zhang 0052, Wassim Hamidouche, Olivier Déforges |
IEEE Trans. Image Process. | 3 |
| 2023 | PAV-SOD: A New Task towards Panoramic Audiovisual Saliency DetectionabstractObject-level audiovisual saliency detection in 360° panoramic real-life dynamic scenes is important for exploring and modeling human perception in immersive environments, also for aiding the development of virtual, augmented, and mixed reality applications in fields such as education, social network, entertainment, and training. To this end, we propose a new task, p anoramic a udio v isual s alient o bject d etection, ( PAV-SOD 1 ), which aims to segment the objects grasping most of the human attention in 360° panoramic videos reflecting real-life daily scenes. To support the task, we collect PAVS10K , the first p anoramic video dataset for a udio v isual s alient object detection, which consists of 67 4K-resolution equirectangular videos with per-video labels including hierarchical scene categories and associated attributes depicting specific challenges for conducting PAV-SOD , and 10,465 uniformly sampled video frames with manually annotated object-level and instance-level pixel-wise masks. The coarse-to-fine annotations enable multi-perspective analysis regarding PAV-SOD modeling. We further systematically benchmark 13 state-of-the-art salient object detection (SOD)/video object segmentation (VOS) methods based on our PAVS10K . Besides, we propose a new baseline network, which takes advantage of both visual and audio cues of 360° video frames by using a new conditional variational auto-encoder (CVAE). Our C VAE-based a udio v isual net work, namely, CAV-Net , consists of a spatial-temporal visual segmentation network, a convolutional audio-encoding network, and audiovisual distribution estimation modules. As a result, our CAV-Net outperforms all competing models and is able to estimate the aleatoric uncertainties within PAVS10K . With extensive experimental results, we gain several findings about PAV-SOD challenges and insights towards PAV-SOD model interpretability. We hope that our work could serve as a starting point for advancing SOD towards immersive media. Yi Zhang 0076, Fang-Yi Chao, Wassim Hamidouche, Olivier Déforges |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Machine Learning Based Efficient Qt-Mtt Partitioning for VVC Inter CodingabstractThe Joint Video Experts Team (JVET) have standardized the Versatile Video Coding (VVC) in 2020 targeting efficient coding of the emerging video services and formats such as 8K and immersive video streaming applications. VVC standard enhances the coding efficiency by 40% at the cost of an encoder computational complexity increase estimated to 859%(x8) compared to the previous standard High Efficiency Video Coding (HEVC). This work aims at reducing the complexity of the VVC encoder under the Random Access (RA) configuration. The proposed method takes advantage of the inter prediction in order to predict the split probabilities through a convolutional neural network. Our solution reaches 31.8% of complexity reduction for a negligible bitrate increase of 1.11% outperforming state-of-the-art methods. Alexandre Tissier, Wassim Hamidouche, Jarno Vanne, Daniel Ménard |
ICIP | 2 |
| 2022 | Channel-Spatial Mutual Attention Network for 360° Salient Object DetectionabstractIn this work, we conduct 360° panoramic salient object detection by taking advantage of both the global and local visual cues of 360° images, with a novel channel-spatial mutual attention network (CSMA-Net). The key component of the CSMA-Net is the proposed CSMA module, which cascades channel-/spatial-weighting-based mutual attentions. The objective of our CSMA module is to refine and fuse the bottleneck features from two separate encoders with different planar representations of 360° panorama as inputs, i.e., equirectangular image and cube map. Our CSMA-Net outperforms 10 state-of-the-art segmentation methods based on the proposed 360° SOD benchmark where multiple fine-tuning and testing strategies are applied to the widely-used 360° datasets. Extensive experimental results illustrate the effectiveness and robustness of the proposed CSMA-Net1. Yi Zhang 0076, Wassim Hamidouche, Olivier Déforges |
ICPR | 2 |
| 2022 | Evaluation of Pre-Trained CNN Models for Geographic Fake Image DetectionabstractThanks to the remarkable advances in generative adversarial networks (GANs), it is becoming increasingly easy to generate/manipulate images. The existing works have mainly focused on deepfake in face images and videos. However, we are currently witnessing the emergence of fake satellite images, which can be misleading or even threatening to national security. Consequently, there is an urgent need to develop detection methods capable of distinguishing between real and fake satellite images. To advance the field, in this paper, we explore the suitability of several convolutional neural network (CNN) architectures for fake satellite image detection. Specifically, we benchmark four CNN models by conducting extensive experiments to evaluate their performance and robustness against various image distortions. This work allows the establishment of new baselines and may be useful for the development of CNN-based methods for fake satellite image detection. Sid Ahmed Fezza, Mohammed Yasser Ouis, Bachir Kaddar, Wassim Hamidouche, Abdenour Hadid |
MMSP | 4 |
| 2022 | Visual Security Evaluation of Perceptually Encrypted Images based on Multi-Task LearningabstractOver past decades, many image encryption algorithms have been proposed, among which we can cite the perceptual/selective encryption methods which have attracted wide attention. Such methods allow for adjusting the scrambling intensity, it is therefore essential to have a reliable visual security metric to adjust the scrambling intensity on the one hand and to evaluate the visual security of encrypted images on the other hand. Usually, these tasks are performed based on classical randomness-based measures or image quality assessment metrics. However, these methods have shown their inadequacy as a visual security metric, as they do not address content intelligibility, which represents an essential security requirement. Moreover, these methods are either dedicated to the prediction of visual security (VS) or visual quality (VQ), but not both. In this paper, we propose a no-reference (NR) visual security metric for perceptually encrypted images based on deep multi-task learning, which we dub the Multi-Task Visual Security (MTVS) metric. The proposed metric consists of one shared convolutional neural network (CNN) followed by two separate sub-networks of fully-connected (FC) layers, where one sub-network is responsible for predicting the VS score, while the other is for predicting the VQ score. Experiments were performed on two publicly perceptually encrypted image databases and the results show that the proposed metric yields superior performance on both VS and VQ prediction tasks. The source code and models are available at: https://github.com/Mamadou-Keita/MTVS. Mamadou Keita, Sid Ahmed Fezza, Wassim Hamidouche, Azeddine Beghdadi |
MMSP | 3 |
| 2022 | Discrete Cosine Basis Oriented Motion Modeling With Cuboidal Applicability Regions For Versatile Video CodingabstractThe relentless expansion of video based applications is underpinned by video coding technologies. The latest video coding standard i.e. versatile video coding (VVC) can provide superior compression performance than its predecessors. In this regard, motion modeling plays a central role. Experimental results showed that the discrete cosine basis oriented motion model can describe complex motion better than an affine motion model, adopted in the VVC. Hence, in this paper we propose to augment the VVC motion modeling technique with a set of discrete cosine basis oriented motion models and the applicability region of each such motion model is determined by non-overlapping rectangular regions, known as cuboids. Experimental results show a bit rate savings of up to 2.37% is achievable with respect to a VVC reference. Ashek Ahmmed, Wassim Hamidouche, Andrew J. Lambert, Mark R. Pickering, M. Manzur Murshed |
PCS | 2 |
| 2022 | Efficient HW Design of Adaptive Loop Filter for 4k ASIC VVC EncoderabstractVersatile Video Coding (VVC) is the next-generation video coding standard released in July 2020. VVC introduces new coding tools enhancing the coding efficiency compared to its predecessor, High Efficiency Video Coding (HEVC). These new tools significantly impact the VVC software and hardware implementations with a complexity estimated to two times and eight times the HEVC decoder and encoder complexity, respectively. In particular, the Adaptive Loop Filter (ALF), adopted in VVC as an in-loop filter, increases both the run time complexity and memory usage. These concerns need to be carefully addressed regarding the design of a VVC hardware encoder. In this paper, we present an efficient hardware implementation of the ALF tool with its decision process in the context of a professional VVC encoder. The proposed solution can reach a real-time encoding of 4K resolution videos with 4:2:2 chroma sub-sampling at 60 frames per second targeting ASIC platforms with 28-nm technology. Ibrahim Farhat, Wassim Hamidouche, Adrien Grill, Daniel Ménard, Olivier Déforges |
PCS | 2 |
| 2022 | Benchmarking Learning-based Bitrate Ladder Prediction Methods for Adaptive Video StreamingabstractHTTP adaptive streaming (HAS) is increasingly adopted by over-the-top (OTT)-based video streaming services, it allows clients to dynamically switch among various stream representations. Each of these representations is encoded to target a specific bitrate providing a wide range of operating bitrates known as the bitrate ladder. Several approaches with different levels of complexity are currently used to build such a bitrate ladder. The most straightforward method is to use a fixed bitrate ladder for all videos, which is a set of bitrate-resolution pairs, called “one-size-fits-all”, and the most complex is based on the intensive encoding of all resolutions over a wide bitrate range to construct the convex-hull. This latter is then used to obtain a per-title bitrate ladder. Recently, various methods relying on machine learning (ML) techniques have been proposed to predict content-based ladder without performing exhaustive search encoding. In this paper, we conduct a benchmark study of several handcrafted and deep learning (DL)-based approaches for predicting content-optimized bitrate ladder, which we believe provides baseline methods and will be useful for future research in this field. The obtained results, based on 200 video sequences compressed with the high-efficiency video coding (HEVC) encoder, reveal that the most efficient method predicts the bitrate ladder without performing any encoding process at the cost of a slight Bjøntegaard delta bitrate (BD-BR) loss of 1.43% compared to the exhaustive approach. The dataset and the source code of the considered methods are made publicly available at: https://github.com/atelili/Bitrate-Ladder-Benchmark. Ahmed Telili, Wassim Hamidouche, Sid Ahmed Fezza, Luce Morin |
PCS | 2 |
| 2022 | Deep multi-task learning for image/video distortions identification
Zoubida Ameur, Sid Ahmed Fezza, Wassim Hamidouche |
Neural Comput. Appl. | 3 |
| 2022 | Detect and defense against adversarial examples in deep learning using natural scene statistics and adaptive denoising
Anouar Kherchouche, Sid Ahmed Fezza, Wassim Hamidouche |
Neural Comput. Appl. | 3 |
| 2022 | Visual Attention-Aware High Dynamic Range Quantization for HEVC Video CodingabstractEmerging HDR videos enable the recording of adequate luminance information and representing realistic scenes to the audience. Due to the high precision of the HDR data recorded in the floating-point format, a quantization process is required to convert HDR data to integer data for compatibility with current transmission and display systems. In this study, a novel attention-aware quantization method is presented that attempts to preserve the contrast details in the region of interest of the human visual system. This method was applied in the context of HDR video coding. The proposed coding solution was compared with the current anchor solution in terms of the quality of the reconstructed video. Experimental results show that the proposed solution is able to improve the visual quality of encoded video with respect to the anchor solution. Additionally, the proposed solution achieves a bit-rate gain over the anchor with reference to the objective evaluation results. Yi Liu 0004, Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Cheolkon Jung |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Learning Synergistic Attention for Light Field Salient Object Detection
Yi Zhang 0076, Geng Chen 0001, Yong Xia 0001, Olivier Déforges, Wassim Hamidouche, Lu Zhang 0037 |
BMVC | 7 |
| 2021 | Multitask Learning for VVC Quality Enhancement and Super-ResolutionabstractThe latest video coding standard, called versatile video coding (VVC), includes several novel and refined coding tools at different levels of the coding chain. These tools bring significant coding gains with respect to the previous standard, high efficiency video coding (HEVC). However, the encoder may still introduce visible coding artifacts, mainly caused by coding decisions applied to adjust the bitrate to the available bandwidth. Hence, pre and post-processing techniques are generally added to the coding pipeline to improve the quality of the decoded video. These methods have recently shown outstanding results compared to traditional approaches, thanks to the recent advances in deep learning. Generally, multiple neural networks are trained independently to perform different tasks, thus omitting to benefit from the redundancy that exists between the models. In this paper, we investigate a learning-based solution as a post-processing step to enhance the decoded VVC video quality. Our method relies on multitask learning to perform both quality enhancement and super-resolution using a single shared network optimized for multiple degradation levels. The proposed solution enables a good performance in both mitigating coding artifacts and super-resolution with fewer network parameters compared to traditional specialized architectures. Charles Bonnineau, Wassim Hamidouche, Jean-François Travers, Naty Ould Sidaty, Olivier Déforges |
PCS | 2 |
| 2021 | Model Selection CNN-based VVC Quality EnhancementabstractArtifact removal and filtering methods are inevitable parts of video coding. On one hand, new codecs and compression standards come with advanced in-loop filters and on the other hand, displays are equipped with high capacity processing units for post-treatment of decoded videos. This paper proposes a Convolutional Neural Network (CNN)-based post-processing algorithm for intra and inter frames of Versatile Video Coding (VVC) coded streams. Depending on the frame type, this method benefits from normative prediction signal by feeding it as an additional input along with reconstructed signal and a Quantization Parameter (QP)-map to the CNN. Moreover, an optional Model Selection (MS) strategy is adopted to pick the best trained model among available ones at the encoder side, and signal it to the decoder side. This MS strategy is applicable at both frame level and block level. The experiments under the Random Access (RA) configuration of the VVC Test Model (VTM-10.0) show that the proposed prediction-aware algorithm can bring an additional BD-BR gain of -1.3% compared to the method without the prediction information. Furthermore, the proposed MS scheme brings -0.5% more BD-BR gain on top of the prediction-aware method. Fatemeh Nasiri, Wassim Hamidouche, Luce Morin, Nicolas Dhollande, Gildas Cocherel |
PCS | 2 |
| 2021 | CAESR: Conditional Autoencoder and Super-Resolution for Learned Spatial ScalabilityabstractIn this paper, we present CAESR, an hybrid learning-based coding approach for spatial scalability based on the versatile video coding (VVC) standard. Our framework considers a low-resolution signal encoded with VVC intra-mode as a base-layer (BL), and a deep conditional autoencoder with hyperprior (AE-HP) as an enhancement-layer (EL) model. The EL encoder takes as inputs both the upscaled BL reconstruction and the original image. Our approach relies on conditional coding that learns the optimal mixture of the source and the upscaled BL image, enabling better performance than residual coding. On the decoder side, a super-resolution (SR) module is used to recover high-resolution details and invert the conditional coding process. Experimental results have shown that our solution is competitive with the VVC full-resolution intra coding while being scalable. Charles Bonnineau, Wassim Hamidouche, Jean-François Travers, Naty Ould Sidaty, Jean-Yves Aubié, Olivier Déforges |
VCIP | 2 |
| 2021 | HCiT: Deepfake Video Detection Using a Hybrid Model of CNN features and Vision TransformerabstractThe number of new falsified video contents is dramatically increasing, making the need to develop effective deepfake detection methods more urgent than ever. Even though many existing deepfake detection approaches show promising results, the majority of them still suffer from a number of critical limitations. In general, poor generalization results have been obtained under unseen or new deepfake generation methods. Consequently, in this paper, we propose a deepfake detection method called HCiT, which combines Convolutional Neural Network (CNN) with Vision Transformer (ViT). The HCiT hybrid architecture exploits the advantages of CNN to extract local information with the ViT's self-attention mechanism to improve the detection accuracy. In this hybrid architecture, the feature maps extracted from the CNN are feed into ViT model that determines whether a specific video is fake or real. Experiments were performed on Faceforensics++ and DeepFake Detection Challenge preview datasets, and the results show that the proposed method significantly outperforms the state-of-the-art methods. In addition, the HCiT method shows a great capacity for generalization on datasets covering various techniques of deepfake generation. The source code is available at: https://github.com/KADDAR-Bachir/HCiT Bachir Kaddar, Sid Ahmed Fezza, Wassim Hamidouche, Zahid Akhtar, Abdenour Hadid |
VCIP | 3 |
| 2021 | Micro-expression recognition from local facial regions
Mouath Aouayeb, Wassim Hamidouche, Catherine Soladié, Kidiyo Kpalma, Renaud Séguier |
Signal Process. Image Commun. | 2 |
| 2021 | Quality-Driven Variable Frame-Rate for Green Video Coding in Broadcast ApplicationsabstractThe Digital Video Broadcasting (DVB) has proposed to introduce the Ultra-High Definition services in three phases: UHD-1 phase 1, UHD-1 phase 2 and UHD-2. The UHD-1 phase 2 specification includes several new features such as High Dynamic Range (HDR) and High Frame-Rate (HFR). It has been shown in several studies that HFR (+100 fps) enhances the perceptual quality and that this quality enhancement is content-dependent. On the other hand, HFR brings several challenges to the transmission chain including codec complexity increase and bit-rate overhead, which may delay or even prevent its deployment in the broadcast echo-system. In this paper, we propose a Variable Frame Rate (VFR) solution to determine the minimum (critical) frame-rate that preserves the perceived video quality of HFR video. The frame-rate determination is modeled as a 3-class classification problem which consists in dynamically and locally selecting one frame-rate among three: 30, 60 and 120 frames per second. Two random forests classifiers are trained with a ground truth carefully built by experts for this purpose. The subjective results conducted on ten HFR video contents, not included in the training set, clearly show the efficiency of the proposed solution enabling to locally determine the lowest possible frame-rate while preserving the quality of the HFR content. Moreover, our VFR solution enables significant bit-rate savings and complexity reductions at both encoder and decoder sides. Glenn Herrou, Charles Bonnineau, Wassim Hamidouche, Patrick Dumenil, Jérôme Fournier, Luce Morin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Light Field Image Coding Using VVC Standard and View Synthesis Based on Dual Discriminator GAN
Nader Bakir, Wassim Hamidouche, Sid Ahmed Fezza, Khouloud Samrouth, Olivier Déforges |
IEEE Trans. Multim. | 2 |
| 2021 | A Multi-FoV Viewport-Based Visual Saliency Model Using Adaptive Weighting Losses for 360$^\circ$ Imagesabstract360$^\circ$media allows observers to explore the scene in all directions. The consequence is that the human visual attention is guided by not only the perceived area in the viewport but also the overall content in 360$^\circ$. In this paper, we propose a method to estimate the 360$^\circ$saliency map which extracts salient features from the entire 360$^\circ$image in each viewport in three different Field of Views (FoVs). Our model is first pretrained with a large-scale 2D image dataset to enable the interpretation of semantic contents, then fine-tuned with a relative small 360$^\circ$image dataset. A novel weighting loss function attached with stretch weighted maps is introduced to adaptively weight the losses of three evaluation metrics and attenuate the impact of stretched regions in equirectangular projection during training process. Experimental results demonstrate that our model achieves better performance with the integration of three FoVs and its diverse viewport images. Results also show that the adaptive weighting losses and stretch weighted maps effectively enhance the evaluation scores compared to the fixed weighting losses solutions. Comparing to other state of the art models, our method surpasses them on three different datasets and ranks the top using 5 performance evaluation metrics on the Salient360! benchmark set. The code is available athttps://github.com/FannyChao/MV-SalGAN360. Fang-Yi Chao, Lu Zhang 0037, Wassim Hamidouche, Olivier Déforges |
IEEE Trans. Multim. | 3 |
| 2020 | Versatile Video Coding and Super-Resolution for Efficient Delivery of 8k Video with 4k Backward-CompatibilityabstractIn this paper, we propose, through an objective study, to compare and evaluate the performance of different coding approaches allowing the delivery of an 8K video signal with 4K backward-compatibility on broadcast networks. Presented approaches include simulcast of 8K and 4K single-layer signals encoded using High-Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC) standards, spatial scalability using SHVC with 4K base layer (BL) and 8K enhancement-layer (EL), and super-resolution applied on 4K VVC signal after decoding to reach 8K resolution. For up-scaling, we selected the deep-learning-based super-resolution method called Super-Resolution with Feedback Network (SRFBN) and the Lanczos interpolation filter. We show that the deep-learning-based approach achieves visual quality gain over simulcast, especially on bit-rates lower than 30Mb/s with average gain of 0.77dB, 0.015, and 7.97 for PSNR, SSIM, and VMAF, respectively and outperforms the Lanczos filter in average by 29% of BD-rate savings. Charles Bonnineau, Wassim Hamidouche, Jean-François Travers, Olivier Déforges |
ICASSP | 2 |
| 2020 | Lightweight Hardware Implementation of VVC Transform Block for ASIC DecoderabstractVersatile Video Coding (VVC) is the next generation video coding standard expected by the end of 2020. Compared to its predecessor, VVC introduces new coding tools to make compression more efficient at the expense of higher computational complexity. This rises a need to design an efficient and optimised implementation especially for embedded platforms with limited memory and logic resources. One of the newly introduced tools in VVC is the Multiple Transform Selection (MTS). This latter involves three Discrete Cosine Transform (DCT)/Discrete Sine Transform (DST) types with larger and rectangular transform blocks. In this paper, an efficient hardware implementation of all DCT/DST transform types and sizes is proposed. The proposed design uses 32 multipliers in a pipelined architecture which targets an ASIC platform. It consists in a multi-standard architecture that supports the transform block of recent MPEG standards including AVC, HEVC and VVC. The architecture is optimized and removes unnecessary complexities found in other proposed architectures by using regular multipliers instead of multiple constant multipliers. The synthesized results show that the proposed method which sustain a constant throughput of two pixels/cycle and constant latency for all block sizes can reach an operational frequency of 600 Mhz enabling to decode in real-time 4K videos at 48 fps. Ibrahim Farhat, Wassim Hamidouche, Adrien Grill, Daniel Ménard, Olivier Déforges |
ICASSP | 2 |
| 2020 | Binary Probability Model for Learning Based Image CompressionabstractIn this paper, we propose to enhance learned image compression systems with a richer probability model for the latent variables. Previous works model the latents with a Gaussian or a Laplace distribution. Inspired by binary arithmetic coding, we propose to signal the latents with three binary values and one integer, with different probability models.A relaxation method is designed to perform gradient-based training. The richer probability model results in a better entropy coding leading to lower rate. Experiments under the Challenge on Learned Image Compression (CLIC) test conditions demonstrate that this method achieves 18 % rate saving compared to Gaussian or Laplace models. Théo Ladune, Pierrick Philippe, Wassim Hamidouche, Lu Zhang 0037, Olivier Déforges |
ICASSP | 3 |
| 2020 | Quality-Driven Dynamic VVC Frame Partitioning for Efficient Parallel ProcessingabstractVVC is the next generation video coding standard, offering coding capability beyond HEVC standard. The high computational complexity of the latest video coding standards requires high-level parallelism techniques, in order to achieve real-time and low latency encoding and decoding. HEVC and VVC include tile grid partitioning that allows to process simultaneously rectangular regions of a frame with independent threads. The tile grid may be further partitioned into a horizontal sub-grid of Rectangular Slices (RSs), increasing the partitioning flexibility. The dynamic Tile and Rectangular Slice (TRS) partitioning solution proposed in this paper benefits from this flexibility. The TRS partitioning is carried-out at the frame level, taking into account both spatial texture of the content and encoding times of previously encoded frames. The proposed solution searches the best partitioning configuration that minimizes the trade-off between multi-thread encoding time and encoding quality loss. Experiments prove that the proposed solution, compared to uniform TRS partitioning, significantly decreases multi-thread encoding time, with slightly better encoding quality. Thomas Amestoy, Wassim Hamidouche, Cyril Bergeron, Daniel Ménard |
ICIP | 2 |
| 2020 | CNN Oriented Complexity Reduction Of VVC Intra EncoderabstractThe Joint Video Expert Team (JVET) is currently developing the next-generation MPEG/ITU video coding standard called Versatile Video Coding (VVC) and their ultimate goal is to double the coding efficiency over the state-of-the-art HEVC standard.The latest version of the VVC reference encoder, VTM6.1, is able to improve the intra coding efficiency by 24 % over the HEVC reference encoder HM16.20, but at the expense of 27 times the encoding time. The complexity overhead of VVC primarily stems from its novel block partitioning scheme that complements Quad-Tree (QT) split with Multi-Type Tree (MTT) partitioning in order to better fit the local variations of the video signal. This work reduces the block partitioning complexity of VTM6.1 through the use of Convolutional Neural Networks (CNNs). For each 64 × 64 Coding Unit (CU), the CNN is trained to predict a probability vector that speeds up coding block partitioning in encoding. Our solution is shown to decrease the intra encoding complexity of VTM6.1 by 51.5% with a bitrate increase of only 1.45%. Alexandre Tissier, Wassim Hamidouche, Jarno Vanne, Franck Galpin, Daniel Ménard |
ICIP | 2 |
| 2020 | A Fixation-Based 360° Benchmark Dataset For Salient Object DetectionabstractFixation prediction (FP) in panoramic contents has been widely investigated along with the booming trend of virtual reality (VR) applications. However, another issue within the field of visual saliency, salient object detection (SOD), has been seldom explored in 360° or omnidirectional) images due to the lack of datasets representative of real scenes with pixel-level annotations. Toward this end, we collect 107 equirectangular panoramas with challenging scenes and multiple object classes. Based on the consistency between FP and explicit saliency judgements, we further manually annotate 1,165 salient objects over the collected images with precise masks under the guidance of real human eye fixation maps. Six state-of-the-art SOD models are then benchmarked on the proposed fixation-based 360° image dataset (F-360iSOD), by applying a multiple cubic projection-based fine-tuning method. Experimental results show a limitation of the current methods when used for SOD in panoramic images, which indicates the proposed dataset is challenging. Key issues for 360° SOD is also discussed. The proposed dataset is available at https://github.com/PanoAsh/F-360iSOD. Yi Zhang 0076, Lu Zhang 0037, Wassim Hamidouche, Olivier Déforges |
ICIP | 3 |
| 2020 | Light Field Image Coding Using Dual Discriminator Generative Adversarial Network And VVC Temporal ScalabilityabstractLight field technology represents a viable path for providing a high-quality VR content. However, such an imaging system generates a high amount of data leading to an urgent need for LF image compression solution. In this paper, we propose an efficient LF image coding scheme based on view synthesis. Instead of transmitting all the LF views, only some of them are coded and transmitted, while the remaining views are dropped. The transmitted views are coded using Versatile Video Coding (VVC) and used as reference views to synthesize the missing views at decoder side. The dropped views are generated using the efficient dual discriminator GAN model. The selection of reference/dropped views is performed using a rate distortion optimization based on the VVC temporal scalability. Experimental results show that the proposed method provides high coding performance and overcomes the state-of-the-art LF image compression solutions. Nader Bakir, Wassim Hamidouche, Sid Ahmed Fezza, Khouloud Samrouth, Olivier Déforges |
ICME | 2 |
| 2020 | Detection of Adversarial Examples in Deep Neural Networks with Natural Scene StatisticsabstractRecent studies have demonstrated that the deep neural networks (DNNs) are vulnerable to carefully-crafted perturbations added to a legitimate input image. Such perturbed images are called adversarial examples (AEs) and can cause DNNs to misclassify. Consequently, it is of paramount importance to develop detection methods of AEs, thus allowing to reject them. In this paper, we propose to characterize the AEs through the use of natural scene statistics (NSS). We demonstrate that these statistical properties are altered by the presence of adversarial perturbations. Based on this finding, we propose three different methods that exploit these scene statistics to determine if an input is adversarial or not. The proposed detection methods have been evaluated against four prominent adversarial attacks and on three standards datasets. The experimental results have shown that the proposed methods achieve a high detection accuracy while providing a low false positive rate. Anouar Kherchouche, Sid Ahmed Fezza, Wassim Hamidouche, Olivier Déforges |
IJCNN | 3 |
| 2020 | Extending 2D Saliency Models for Head Movement Prediction in 360-Degree Images using CNN-Based FusionabstractSaliency prediction can be of great benefit for 360-degree image/video applications, including compression, streaming, rendering and viewpoint guidance. It is therefore quite natural to adapt the 2D saliency prediction methods for 360-degree images. To achieve this, it is necessary to project the 360-degree image to 2D plane. However, the existing projection techniques introduce different distortions, which provides poor results and makes inefficient the direct application of 2D saliency prediction models to 360-degree content. Consequently, in this paper, we propose a new framework for effectively applying any 2D saliency prediction method to 360-degree images. The proposed framework particularly includes a novel convolutional neural network based fusion approach that provides more accurate saliency prediction while avoiding the introduction of distortions. The proposed framework has been evaluated with five 2D saliency prediction methods, and the experimental results showed the superiority of our approach compared to the use of weighted sum or pixel-wise maximum fusion methods. Ibrahim Djemai, Sid Ahmed Fezza, Wassim Hamidouche, Olivier Déforges |
ISCAS | 3 |
| 2020 | Natural Scene Statistics for Detecting Adversarial Examples in Deep Neural NetworksabstractThe deep neural networks (DNNs) have been adopted in a wide spectrum of applications. However, it has been demonstrated that their are vulnerable to adversarial examples (AEs): carefully-crafted perturbations added to a clean input image. These AEs fool the DNNs which classify them incorrectly. Therefore, it is imperative to develop a detection method of AEs allowing the defense of DNNs. In this paper, we propose to characterize the adversarial perturbations through the use of natural scene statistics. We demonstrate that these statistical properties are altered by the presence of adversarial perturbations. Based on this finding, we design a classifier that exploits these scene statistics to determine if an input is adversarial or not. The proposed method has been evaluated against four prominent adversarial attacks and on three standards datasets. The experimental results have shown that the proposed detection method achieves a high detection accuracy, even against strong attacks, while providing a low false positive rate. Anouar Kherchouche, Sid Ahmed Fezza, Wassim Hamidouche, Olivier Déforges |
MMSP | 3 |
| 2020 | Optical Flow and Mode Selection for Learning-based Video CodingabstractThis paper introduces a new method for inter-frame coding based on two complementary autoencoders: MOFNet and CodecNet. MOFNet aims at computing and conveying the Optical Flow and a pixel-wise coding Mode selection. The optical flow is used to perform a prediction of the frame to code. The coding mode selection enables competition between direct copy of the prediction or transmission through CodecNet.The proposed coding scheme is assessed under the Challenge on Learned Image Compression 2020 (CLIC20) P-frame coding conditions, where it is shown to perform on par with the state-of-the-art video codec ITU/MPEG HEVC. Moreover, the possibility of copying the prediction enables to learn the optical flow in an end-to-end fashion i.e. without relying on pre-training and/or a dedicated loss term. Théo Ladune, Pierrick Philippe, Wassim Hamidouche, Lu Zhang 0037, Olivier Déforges |
MMSP | 3 |
| 2020 | Towards Audio-Visual Saliency Prediction for Omnidirectional Video with Spatial AudioabstractOmnidirectional videos (ODVs) with spatial audio enable viewers to perceive 360° directions of audio and visual signals during the consumption of ODVs with head-mounted displays (HMDs). By predicting salient audio-visual regions, ODV systems can be optimized to provide an immersive sensation of audio-visual stimuli with high-quality. Despite the intense recent effort for ODV saliency prediction, the current literature still does not consider the impact of auditory information in ODVs. In this work, we propose an audio-visual saliency (AVS360) model that incorporates 360° spatial-temporal visual representation and spatial auditory information in ODVs. The proposed AVS360 model is composed of two 3D residual networks (ResNets) to encode visual and audio cues. The first one is embedded with a spherical representation technique to extract 360° visual features, and the second one extracts the features of audio using the log mel-spectrogram. We emphasize sound source locations by integrating audio energy map (AEM) generated from spatial audio description (i.e., ambisonics) and equator viewing behavior with equator center bias (ECB). The audio and visual features are combined and fused with AEM and ECB via attention mechanism. Our experimental results show that the AVS360 model has significant superiority over five state-of-the-art saliency models. To the best of our knowledge, it is the first w ork that develops the audio-visual saliency model in ODVs. The code will be publicly available to foster future research on audio-visual saliency in ODVs. Fang-Yi Chao, Cagri Ozcinar, Lu Zhang 0037, Wassim Hamidouche, Olivier Déforges, Aljoscha Smolic |
VCIP | 4 |
| 2020 | Prediction-Aware Quality Enhancement of VVC Using CNNabstractThe upcoming video coding standard, Versatile Video Coding (VVC), has shown great improvement compared to its predecessor, High Efficiency Video Coding (HEVC), in terms of bitrate saving. Despite its substantial performance, compressed videos might still suffer from quality degradation at low bitrates due to coding artifacts such as blockiness, blurriness and ringing. In this work, we exploit Convolutional Neural Networks (CNN) to enhance quality of VVC coded frames after decoding in order to reduce low bitrate artifacts. The main contribution of this work is the use of coding information from the compressed bitstream. More precisely, the prediction information of intra frames is used for training the network in addition to the reconstruction information. The proposed method is applied on both luminance and chrominance components of intra coded frames of VVC. Experiments on VVC Test Model (VTM) show that, both in low and high bitrates, the use of coding information can improve the BD-rate performance by about 1% and 6% for luma and chroma components, respectively. Fatemeh Nasiri, Wassim Hamidouche, Luce Morin, Nicolas Dhollande, Gildas Cocherel |
VCIP | 2 |
| 2020 | Software HEVC video decoder: towards an energy saving for mobile applications
Naty Ould Sidaty, Julien Heulot, Wassim Hamidouche, Maxime Pelcat, Daniel Ménard |
Multim. Tools Appl. | 3 |
| 2020 | Forward-Inverse 2D Hardware Implementation of Approximate Transform Core for the VVC StandardabstractThe future video coding standard named Versatile Video Coding (VVC) is expected by the end of 2020. VVC will enable better coding efficiency than the current High Efficiency Video Coding (HEVC) standard. This coding gain is brought by several coding tools. The Multiple Transform Selection (MTS) is one of the key coding tools that have been introduced in VVC. The MTS concept relies on three transform types including Discrete Cosine Transform (DCT)-II, Discrete Sine Transform (DST)-VII and DCT-VIII. Unlike the DCT-II that has fast computing algorithms, the DST-VII and DCT-VIII rely on more complex matrix multiplication. In this paper an approximation approach is proposed to reduce the computational cost of the DST-VII and DCT-VIII. The approximation consists in applying adjustment stages, based on sparse block-band matrices, to a variant of DCT-II family mainly DCT-II and its inverse. Genetic algorithm is used to derive the optimal coefficients of the adjustment matrices. Moreover, an efficient hardware implementation of the forward and inverse approximate transform module is proposed. The architecture design includes a pipelined and reconfigurable forward-inverse DCT-II core transform as it is the main core for DST-VII and DCT-VIII computations. The proposed 32-point 1D architecture including low cost adjustment stages allows the processing of a video in 2K and 4K resolutions at 1095 and 273 frames per second, respectively. A unified 2D implementation of forward-inverse DCT-II, approximate DST-VII and DCT-VIII is also presented. The synthesis results show that the design is able to sustain a video in 2K and 4K resolutions at 386 and 96 frames per second, respectively, while using only 12% of Alms, 22% of registers and 30% of DSP blocks of the Arria10 SoC platform. Ahmed Kammoun, Wassim Hamidouche, Pierrick Philippe, Olivier Déforges, Fatma Belghith, Nouri Masmoudi, Jean-François Nezan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Tunable VVC Frame Partitioning Based on Lightweight Machine LearningabstractBlock partition structure is a critical module in video coding scheme to achieve significant gap of compression performance. Under the exploration of the future video coding standard, named Versatile Video Coding (VVC), a new Quad Tree Binary Tree (QTBT) block partition structure has been introduced. In addition to the QT block partitioning defined in High Efficiency Video Coding (HEVC) standard, new horizontal and vertical BT partitions are enabled, which drastically increases the encoding time compared to HEVC. In this paper, we propose a lightweight and tunable QTBT partitioning scheme based on a Machine Learning (ML) approach. The proposed solution uses Random Forest classifiers to determine for each coding block the most probable partition modes. To minimize the encoding loss induced by misclassification, risk intervals for classifier decisions are introduced in the proposed solution. By varying the size of risk intervals, tunable trade-off between encoding complexity reduction and coding loss is achieved. The proposed solution implemented in the JEM-7.0 software offers encoding complexity reductions ranging from 30average for only 0.7% to 3.0% Bjxntegaard Delta Rate (BDBR) increase in Random Access (RA) coding configuration, with very slight overhead induced by Random Forest. The proposed solution based on Random Forest classifiers is also efficient to reduce the complexity of the Multi-Type Tree (MTT) partitioning scheme under the VTM-5.0 software, with complexity reductions ranging from 25% to 61% in average for only 0.4% to 2.2% BD-BR increase. Thomas Amestoy, Alexandre Mercat, Wassim Hamidouche, Daniel Ménard, Cyril Bergeron |
IEEE Trans. Image Process. | 3 |
| 2019 | RDO-Based Light Field Image Coding Using Convolutional Neural Networks and Linear ApproximationabstractThe increasing penetration of acquisition and display devices for Light Field (LF) content in the consumer market leads to the high proliferation of this new immersive media. This growing interest to LF images thus urgently raises the question of their compression. In this paper, we propose a convolutional neural networks (CNN)-based LF image coding scheme including both Rate Distortion Optimization (RDO) and post-processing steps. First, at the encoder side, the views are rearranged in sparse and dropped set of views. The former are compressed with a standard encoder and transmitted, while the dropped views are either linearly approximated or synthesized by a CNN using the encoded views as input. This choice is made on the basis of the proposed RDO process. At the decoder side, once the dropped views are either linearly approximated or synthesized by a CNN block, a post-processing step is performed to further enhance the quality of the reconstructed views. This post-processing block is based on superpixel to pixel-matching. Experimental results show that the proposed scheme provides views with high visual quality and overcomes the state-of-the-art LF image compression solutions by -30% in terms of BD-BR and 0.62 dB in BD-PSNR. Nader Bakir, Wassim Hamidouche, Olivier Déforges, Khouloud Samrouth, Sid Ahmed Fezza |
DCC | 2 |
| 2019 | Dynamic Lists for Efficient Coding of Intra Prediction Modes in the Future Video Coding StandardabstractThe next generation MPEG video coding standard is under development by the Joint Video Coding Experts Team (JVET). This new standard, called Versatile Video Coding (VVC), is expected by the end of 2020 and will offer better coding efficiency than its predecessor High Efficiency Video Coding (HEVC) standard. This coding gain is enabled by new coding tools such as more flexible block partitioning, more accurate Intra/Inter predictions, multiple transforms and adaptive in-loop filtering. In this paper we focus on the coding of the Intra Prediction Modes (IPM) that have been increased from 35 modes in HEVC to 67 modes in VVC. We propose a solution based on genetic algorithms to build an ordered list for the coding of IPM in the Joint Exploration Model (JEM) codec. We first give the theoretical upper bound performance in terms of required bits per IPM to encode the IPM using the available contextual information. The new ordering of the labels associated with more efficient codes is then proposed to efficiently leverage contextual informations available in the encoder and construct the Most Probable Modes (MPM) list. The proposed coding scheme enables to increase the BD-BR performance in average by 0.09% for the same level of complexity compared to the JEM. Kevin Reuze, Wassim Hamidouche, Pierrick Philippe, Olivier Déforges |
DCC | 2 |
| 2019 | Random Forest Oriented Fast QTBT Frame PartitioningabstractBlock partition structure is a critical module in video coding scheme to achieve significant gap of compression performance. Under the exploration of future video coding standard by the Joint Video Exploration Team (JVET), named Versatile Video Coding (VVC), a new Quad Tree Binary Tree (QTBT) block partition structure has been introduced. In addition to the QT block partitioning defined by High Efficiency Video Coding (HEVC) standard, new horizontal and vertical BT partitions are enabled, which drastically increases the encoding time compared to HEVC. In this paper, we propose a fast QTBT partitioning scheme based on a Machine Learning approach. Complementary to techniques proposed in literature to reduce the complexity of HEVC Quad Tree (QT) partitioning, the propose solution uses Random Forest classifiers to determine for each block which partition modes between QT and BT is more likely to be selected. Using uncertainty zones of classifier decisions, the proposed complexity reduction technique is able to reduce in average by 30% the encoding time of JEM-v7.0 software in Random Access configuration with only 0.57% Bjøntegaard Delta Rate (BD-BR) increase. Thomas Amestoy, Alexandre Mercat, Wassim Hamidouche, Cyril Bergeron, Daniel Ménard |
ICASSP | 3 |
| 2019 | Rate-Distortion Optimized Tree-Structured Point-Lattice Vector Quantization for Compression of 3D Point Clouds GeometryabstractThis paper deals with the current trends of new compression methods for 3-D point cloud contents required to ensure efficient transmission and storage. The representation of 3D point clouds geometry remains a challenging problem, since this signal is unstructured. In this paper, we introduce a new hierarchical geometry representation based on adaptive Tree-Structured Point-Lattice Vector Quantization (TSPLVQ). This representation enables hierarchically structured 3D content that improves the compression performance for static point clouds. The novelty of the proposed scheme lies in adaptive selection of the optimal quantization scheme of the geometric information, that better leverage the intrinsic correlations in point cloud. Based on its adaptive and multiscale structure, two quantization schemes are dedicated to project recursively the 3D point clouds into a series of embedded truncated cubic lattices. At each step of the process, the optimal quantization scheme is selected according to a rate-distortion cost in order to achieve the best trade-off between coding rate and geometry distortion, such that the compression flexibility and performance can be greatly improved. Experimental results show the interest of the proposed multi-scale method for lossy compression of geometry. Amira Filali, Vincent Ricordel, Nicolas Normand, Wassim Hamidouche |
ICIP | 4 |
| 2019 | Low-Complexity Scalable Encoder Based on Local Adaptation of the Spatial ResolutionabstractA two-layer low-complexity scalable encoding scheme based on local Adaptive Spatial Resolution (ASR) is proposed. This scheme relies on a block-level spatial resolution adaptation in the enhancement layer encoder. For each block, the optimal resolution is either obtained via a rate-distortion optimization followed by a decision refinement process or by a prediction via motion compensation exploiting the base layer motion vectors. The proposed architecture has been integrated over of the High Efficiency Video Coding (HEVC) reference software (HM16.12) which is used as a base layer encoder. Compared to SHVC, the scalable extension of HEVC, experimental results show bitrate savings of 0.76 % as well as encoding complexity reductions of 47 % for the whole scalable encoder and 96 % for the enhancement layer encoder. Glenn Herrou, Wassim Hamidouche, Luce Morin |
ICIP | 2 |
| 2019 | Complexity Reduction Opportunities in the Future VVC Intra EncoderabstractThe Joint Video Expert Team (JVET) is developing the next-generation video coding standard called Versatile Video Coding (VVC) and their ultimate goal is to double the coding efficiency over the current state-of-the-art standard HEVC without letting complexity get out of hand. This work addresses the complexity of the VVC reference encoder called VVC Test Model (VTM) under All Intra coding configuration. The VTM3.0 is able to improve intra coding efficiency by 21% over the latest HEVC reference encoder HM16.19. This coding gain primarily stems from three new coding tools. First, the HEVC Quad-Tree (QT) structure extension with Multi-Type Tree (MTT) partitioning. Second, the duplication of intra prediction modes from 35 to 67. And third, the Multiple Transform Selection (MTS) scheme with two new discrete cosine/sine transforms (DCT-VIII and DST-VII). However, these new tools also play an integral part in making VTM intra encoding around 20 times as complex as that of HM. The purpose of this work is to analyze these tools individually and specify theoretical upper limits for their complexity reduction. According to our evaluations, the complexity reduction opportunity of block partitioning is up to 97%, i.e., the encoding complexity would drop down to 3% for the same coding efficiency if the optimal block partitioning could be directly predicted. The respective percentages for intra mode reduction and MTS optimization are 65% and 55%. We believe these results motivate VVC codec designers to develop techniques that are able to take most out of these opportunities. Alexandre Tissier, Alexandre Mercat, Thomas Amestoy, Wassim Hamidouche, Jarno Vanne, Daniel Ménard |
MMSP | 4 |
| 2019 | Hardware-friendly DST-VII/DCT-VIII approximations for the Versatile Video Coding StandardabstractVersatile Video Coding (VVC) is the next generation video coding standard expected by the end of 2020. The new concept of Multiple-Transform Selection (MTS) has been introduced in VVC. MTS enables the VVC encoder to select the transform that minimizes the rate-distortion cost among a set of pre-defined trigonometric transforms including the well known Discrete Cosine Transform (DCT)-II, DCT-VIII and Discrete Sine Transform (DST)-VII. Unlike the DCT-II that has fast computing algorithms, the DST-VII and DCT-VIII rely on more complex matrix multiplication.This paper tackles the problem of DST-VII and DCT-VIII approximations based on the DCT-II and an adjustment stage. This latter consists in a multiplication by a band-matrix with low number of non-zero coefficients per row. The approximation problem is first modeled as a constrained integer optimization problem minimizing both error and orthogonality. The genetic algorithm is then used to solve the optimization problem and find the adjustment band-matrix that minimizes a trade-off between error and orthogonality. The proposed solution enables to preserve the coding gain achieved by the MTS and considerably reduces the complexity in terms of required number of multiplications by coefficient. Moreover, the proposed approach is hardwarefriendly and will provide a lightweight shared hardware module for DST-II, DST-VII and DCT-VIII transforms. Wassim Hamidouche, Pierrick Philippe, Camar-Eddine Mohamed, Ahmed Kammoun, Daniel Ménard, Olivier Déforges |
PCS | 1 |
| 2019 | Compression Performance of the Versatile Video Coding: HD and UHD Visual Quality MonitoringabstractVideo compression and content quality have become one of the most research topic in the recent years. Predominantly, trends obviously signpost that the video usage over the Internet is on the upsurge. Simultaneously, users' requirement for enlarged resolution and higher quality is rising. Consequently, a huge effort has been made for video coding technologies and quality monitoring. In this paper, we present a subjective-based comparison as well as an objective measurement between the newest Versatile Video Coding (VVC) and the well-known High Efficiency Video Coding (HEVC) standards. Several videos of various content are selected as tested sequences. Both High Definition (HD) and Ultra High Definition (UHD) resolutions are used in this experiment. An extensive range of bit-rates from low to high bit-rates were selected. These sequences are encoded using both HEVC reference software (HM-16.2) and the latest reference software of VVC (VTM-5.0). Obtained results have shown that VVC outperforms consistently HEVC, for realistic bit rates and quality levels, in the range of 40% on the subjective scale. For the objective measurements, using PSNR, SSIM and VMAF as quality metrics, the quality enhancement of VVC over HEVC is ranging from 31% to 40%, depending on video content and spatial resolution. Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Pierrick Philippe, Jérôme Fournier |
PCS | 2 |
| 2019 | Perceptual Evaluation of Adversarial Attacks for CNN-based Image ClassificationabstractDeep neural networks (DNNs) have recently achieved state-of-the-art performance and provide significant progress in many machine learning tasks, such as image classification, speech processing, natural language processing, etc. However, recent studies have shown that DNNs are vulnerable to adversarial attacks. For instance, in the image classification domain, adding small imperceptible perturbations to the input image is sufficient to fool the DNN and to cause misclassification. The perturbed image, called adversarial example, should be visually as close as possible to the original image. However, all the works proposed in the literature for generating adversarial examples have used the Lpnorms (L0, L2and L∞) as distance metrics to quantify the similarity between the original image and the adversarial example. Nonetheless, the Lpnorms do not correlate with human judgment, making them not suitable to reliably assess the perceptual similarity/fidelity of adversarial examples. In this paper, we present a database for visual fidelity assessment of adversarial examples. We describe the creation of the database and evaluate the performance of fifteen state-of-the-art full-reference (FR) image fidelity assessment metrics that could substitute Lpnorms. The database as well as subjective scores are publicly available to help designing new metrics for adversarial examples and to facilitate future research works. Sid Ahmed Fezza, Yassine Bakhti, Wassim Hamidouche, Olivier Déforges |
QoMEX | 3 |
| 2019 | Visual Security Assessment of Selective Video EncryptionabstractGiven the wide use of videos in various applications and across different devices, this raises the question of their security and confidentiality. In the last decade, many video encryption methods have been proposed in the literature. Accordingly, it becomes necessary to have a reliable assessment tool allowing evaluation of the efficiency of these video encryption methods, especially from the visual security point of view. Usually, the visual security is evaluated through the classical objective signal-based metrics. However, these metrics showed their limits as visual security metric, since they are not designed to deal with the security requirements, such as the determination of content intelligibility. Despite its obvious importance, very few visual security metrics have been proposed for the assessment of video encryption methods. This is mainly due to the lack of ground truth with subjective human scores for video encryption applications. In this paper, we present a new database for visual security assessment of selective video encryption. The database including unencrypted and encrypted video contents generated using different selective encryption schemes, as well as subjective scores, is publicly available to help designing new visual security metrics1. Sid Ahmed Fezza, Wassim Hamidouche, Reda Abdellah Kamraoui, Olivier Déforges |
QoMEX | 2 |
| 2019 | An Adaptive Quantizer for High Dynamic Range Content: Application to Video CodingabstractIn this paper, we propose an adaptive perceptual quantization method to convert the representation of high dynamic range (HDR) content from the floating point data type to integer, which is compatible with the current image/video coding and display systems. The proposed method considers the luminance distribution of the HDR content, as well as the detectable contrast threshold of the human visual system, in order to preserve more contrast information than the perceptual quantizer (PQ) in integer representation. Aiming to demonstrate the effectiveness of this quantizer for HDR video compression, we implemented it in a mapping function on the top of the HDR video coding system based on high efficiency video coding standard. Moreover, a comparison function is also introduced to decrease the additional bit-rate of side information, generated by the mapping function. Objective quality measurements and subjective tests have been conducted in order to evaluate the quality of the reconstructed HDR videos. Subjective test results have shown that the proposed method can improve, in a significant manner, the perceived quality of some reconstructed HDR videos. In the objective assessment, the proposed method achieves improvements over PQ in terms of the average bit-rate gain for metrics used in the measurement. Yi Liu 0004, Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Giuseppe Valenzise, Emin Zerman |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Optimal Adaptive Quantization Based on Temporal Distortion Propagation Model for HEVCabstractOptimal adaptive quantization is one of the key points to optimize the coding efficiency of video encoders. The latest block-based video compression standards, such as high-efficiency video coding (HEVC), extensively use predictive coding techniques that create dependencies between blocks and increase the complexity of optimal block quantizers search. Specifically, the motion compensation is responsible for a dependency network connecting all blocks of the same GOP together. In this paper, this dependency network is estimated by a temporal distortion propagation model and an accurate estimation of Inter and Skip modes probabilities. Optimal quantizers are then designed per block in order to achieve global optimization in terms of rate-distortion efficiency. By implementing the algorithm into the HEVC reference model (HM), we report -16.51% PSNR-based and -26.26% SSIM-based average bitrate savings compared to no adaptive quantization. The proposed algorithm outperforms several related methods from the state-of-the-art. Moreover, along with the demonstration of an optimal quantizer solution, we propose an in-depth analysis of the algorithm behavior. This analysis includes, among others, the relative distribution of rates between frames and the control of quantizers dynamic range. Maxime Bichon, Julien Le Tanou, Michaël Ropert, Wassim Hamidouche, Luce Morin |
IEEE Trans. Image Process. | 4 |
| 2018 | Low-Complexity Spatial Scalability Scheme Using HEVC for 4K and VR VideosabstractScalable video coding enables to compress the video at different formats within a single layered bitstream. SHVC, the scalable extension of the High Efficiency Video Coding (HEVC) standard, enables x2 spatial scalability, among other additional features. The closed-loop architecture of the SHVC codec is based on the use of multiple instances of the HEVC codec to encode the video layers, which considerably increases the encoding complexity. With the arrival of new immersive video formats, like 4K, 8K, High Frame Rate (HFR) and 360° videos, the quantity of data to compress is exploding, making the use of high-complexity coding algorithms unsuitable. In this paper, we propose a low-complexity scalable coding scheme based on the use of a single HEVC codec instance and a wavelet-based decomposition as preprocessing. The pre-encoding image decomposition relies on well-known simple Discrete Wavelet Transform (DWT) kernels, such as Haar or Le Gall 5/3. Compared to SHVC, the proposed architecture achieves a similar rate distortion performance with a coding complexity reduction of 50%. Glenn Herrou, Wassim Hamidouche, Luce Morin |
DCC | 2 |
| 2018 | Low Complexity Joint RDO of Prediction Units Couples for HEVC Intra CodingabstractHEVC is the latest block-based video compression standard, outperforming H.264/AVC by 50% bitrate savings for the same perceptual quality. An HEVC encoder provides Rate-Distortion optimization coding tools for block-wise compression. Because of complexity limitations, Rate-Distortion Optimization (RDO) is usually performed independently for each block, assuming coding efficiency losses to be negligible. In this paper, we propose an acceleration solution for the Intra coding scheme named Dual-JRDO, which takes advantage of Inter-Block dependencies related to both predictive coding and CABAC. The Dual-JRDO improves Intra coding efficiency at the expense of higher computational complexity. The acceleration of the Dual-JRDO scheme includes adaptive use of the Dual-JRDO model based on source analysis, short-listing and early decisions strategies. The proposed Fast Dual-JRDO reduces the original model complexity by 89.54%, while providing tractable computation for average R-D gains of -0.45% (up to -0.82%) in the HM16.12 reference software model. Maxime Bichon, Julien Le Tanou, Michaël Ropert, Wassim Hamidouche, Luce Morin, Lu Zhang 0037 |
ICASSP | 4 |
| 2018 | Light Field Image Compression Based on Convolutional Neural Networks and Linear ApproximationabstractComputer vision applications such as refocusing, segmentation and classification become one of the most advanced imaging services. Light Field (LF) imaging systems provide a rich semantic information of the scene. Using a dense set of cameras and microlens arrays (Plenoptic camera), the direction of each ray coming from the scene toward the LF capture system can be extracted and represented by spatial and angular coordinates. However, such imaging system induces many drawbacks including the large amount of data produced and complexity increase for scene representation. In this paper, we propose an efficient LF image coding scheme. This scheme first encodes a sparse set of views using the latest hybrid video encoder (JEM). Then, it estimates a second sparse set of views using a linear approximation. At the decoder side, we use a Deep Learning (DL) approach to estimate the whole LF image from the reconstructed sparse sets of views. Experimental results show that the proposed scheme provides higher visual quality and overcomes the state of the art LF image compression solution by 30 % bitrate gain. Nader Bakir, Wassim Hamidouche, Olivier Déforges, Khouloud Samrouth |
ICIP | 2 |
| 2018 | Live Demonstration: End-to-End Real-Time ROI-based Encryption in HEVC VideosabstractThis paper presents a demonstration setup for live HEVC video coding with Region of Interest (ROI) encryption. The showcased approach splits video frames into independent HEVC tiles and encrypts those belonging to the ROI. This end-to-end content protection scheme is put into practice by integrating the algorithms of selective encryption into Kvazaar HEVC encoder and decryption into openHEVC decoder. The shown implementation performs secure encryption of the ROI in real time with small bit rate and complexity overhead. Naty Ould Sidaty, Marko Viitanen, Wassim Hamidouche, Jarno Vanne, Olivier Déforges |
ISCAS | 3 |
| 2018 | Backward Compatible Layered Video Coding for 360° Video BroadcastabstractRecently, coding of 360° video contents has been investigated in the context of over-the-top streaming services. To be delivered using terrestrial broadcast, it is required to provide backward compatibility of such content to legacy receivers. In this paper, a novel layered coding scheme is proposed to address the delivery of 360° video content over terrestrial broadcast networks. One or several views are extracted from the 360° video and coded as base layers using standard HEVC encoding. Inter-layer reference pictures are built based on projected base-layers and are used in the enhancement layer to encode the 360° video. Experimental results show that the proposed approach provides substantial coding gains of 14.99% compared to simulcast coding and enables limited coding overhead of 5.15% compared to 360° single-layer coding. Thibaud Biatek, Jean-François Travers, Pierre-Loup Cabarat, Wassim Hamidouche |
PCS | 4 |
| 2018 | Temporal Adaptive Quantization using Accurate Estimations of Inter and Skip ProbabilitiesabstractHybrid video coding systems use spatial and temporal predictions in order to remove redundancies within the video source signal. These predictions create coding-scheme-related dependencies, often neglected for sake of simplicity. The R-D Spatio-Temporal Adaptive Quantization (RDSTQ) solution uses such dependencies to achieve better coding efficiency. It models the temporal distortion propagation by estimating the probability of a Coding Unit (CU) to be Inter coded. Uased on this probability, each CU is given a weight depending on its relative importance compared to other CUs. However, the initial approach roughly estimates the Inter probability and does not take into account the Skip mode characteristics in the propagation. It induces important Target uitrate Deviation (TBD) compared to the reference target rate. This paper provides undeniable improvements of the original RDSTQ model in using a more accurate estimation of the Inter probability. Then a new analytical solution for local quantizers is obtained by introducing the Skip probability of a CU into the temporal distortion propagation model. The proposed solution brings -2.05% BD-BR gain in average over the RDSTQ at low rate, which corresponds to -13.54% BD-BR gain in average against no local quantization. Moreover, the TBD is reduced from 38% to 14%. Maxime Bichon, Julien Le Tanou, Michaël Ropert, Wassim Hamidouche, Luce Morin, Lu Zhang 0037 |
PCS | 4 |
| 2018 | Wavelet Decomposition Pre-processing for Spatial Scalability Video Compression SchemeabstractScalable video coding enables to compress the video at different formats within a single layered bitstream. SHVC, the scalable extension of the High Efficiency Video Coding (HEVC) standard, enables x2 spatial scalability, among other additional features. The closed-loop architecture of the SHVC codec is based on the use of multiple instances of the HEVC codec to encode the video layers, which considerably increases the encoding complexity. With the arrival of new immersive video formats, like 4K, 8K, High Frame Rate (HFR) and 360° videos, the quantity of data to compress is exploding, making the use of high-complexity coding algorithms unsuitable. In this paper, we propose a lowcomplexity scalable coding scheme based on the use of a single HEVC codec instance and a wavelet-based decomposition as pre-processing. The pre-encoding image decomposition relies on well-known simple Discrete Wavelet Transform (DWT) kernels, such as Haar or Le Gall 5/3. Compared to SHVC, the proposed architecture achieves a similar rate distortion performance with a coding complexity reduction of 50%. Glenn Herrou, Wassim Hamidouche, Luce Morin |
PCS | 2 |
| 2018 | Machine Learning Based Choice of Characteristics for the One-Shot Determination of the HEVC Intra Coding TreeabstractIn the last few years, the Internet of Things (IoT) has become a reality. Forthcoming applications are likely to boost mobile video demand to an unprecedented level. A large number of systems are likely to integrate the latest MPEG video standard High Efficiency Video Coding (HEVC) in the long run and will particularly require energy efficiency. In this context, constraining the computational complexity of embedded HEVC encoders is a challenging task, especially in the case of software encoders. The most energy consuming part of a software intra encoder is the determination of the coding tree partitioning, i.e. the size of pixel blocks. This determination usually requires an iterative process that leads to repeating some encoding tasks. State-of-the-art studies have focused on predicting, from “easily” computed characteristics, an efficient coding tree. They have proposed and evaluated independently many characteristics for one-shot quad-tree prediction. In this paper, we present a fair comparison of these characteristics using a Machine Learning approach and a real-time HEVC encoder. Both computational complexity and information gain are considered, showing that characteristics are far from equivalent in terms of coding tree prediction performance. Alexandre Mercat, Florian Arrestier, Maxime Pelcat, Wassim Hamidouche, Daniel Ménard |
PCS | 4 |
| 2018 | Evaluation of No-reference quality metrics for Ultrasound liver imagesabstractAlthough assessing post-processed medical images is still done by radiologists (rather than computers), numerous algorithms dedicated to medical image processing are developed without taking into consideration the expert's perceived quality scores. In order to evaluate these algorithms, we study in this paper four No-Reference(NR) quality assessment metrics in terms of correlation with perceived scores of experts. These scores were obtained through subjective tests conducted on ultrasound (US) livers images. Results show that one NR metric among the four evaluated performs the best for assessing the quality of US images. However, further study is needed for the development of more suitable NR metrics. Meriem Outtas, Lu Zhang 0037, Olivier Déforges, Wassim Hamidouche, Amina Serir |
QoMEX | 4 |
| 2018 | Reproducible Evaluation of System Efficiency With a Model of Architecture: From Theory to PracticeabstractCurrent trends in high performance and embedded computing include design of increasingly complex hardware architectures with high parallelism, heterogeneous processing elements, and nonuniform communication resources. In order to take hardware and software design decisions, early evaluations of the system nonfunctional properties are needed. These evaluations of system efficiency require electronic system-level information on both algorithms and architecture. Contrary to algorithm models for which a major body of work has been conducted on defining formal models of computation (MoCs), architecture models from the literature are mostly empirical models from which reproducible experimentation requires the accompanying software. In this paper, a precise definition of a model of architecture (MoA) is proposed that focuses on reproducibility and abstraction and removes the overlap previously existing between the notions of MoA and MoC. A first MoA, called the linear system-level architecture model (LSLA), is presented. To demonstrate the generic nature of the proposed new architecture modeling concepts, we show that the LSLA model can be integrated flexibly with different MoCs. LSLA is then used to model the energy consumption of a state-of-the-art multiprocessor system-on-chip (MPSoC) when running an application described using the synchronous dataflow MoC. A method to automatically learn LSLA model parameters from platform measurements is introduced. Despite the high complexity of the underlying hardware and software, a simple LSLA model is demonstrated to estimate the energy consumption of the MPSoC with a fidelity of 86%. Maxime Pelcat, Alexandre Mercat, Karol Desnos, Luca Maggiani, Yanzhou Liu 0001, Julien Heulot, Jean-François Nezan, Wassim Hamidouche, Daniel Ménard, Shuvra S. Bhattacharyya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2017 | Exploiting computation skip to reduce energy consumption by approximate computing, an HEVC encoder case studyabstractApproximate computing paradigm provides methods to optimize algorithms with considering both computational accuracy and complexity. This paradigm can be exploited at different levels of abstraction, from technological to application levels. Approximate computing at algorithm level aims at reducing computational complexity by approximating or skipping block functions of the computation. Numerous applications in the signal and image processing domain integrate algorithms based on discrete optimization techniques. These techniques minimize a cost function by exploring the search space. In this paper, a new approach is proposed to exploit the computation-skipping approximate computing concept by using the Smart Search Space Reduction (Sssr) technique. Sssr enables early selection of the best candidate configurations to reduce the search space. An efficient SSSR technique adjusts configuration selectivity to reduce execution complexity while selecting the most suitable functions to skip. The High Efficiency Video Coding (HEVC) encoder in All Intra (AI) profile is used as a case study to illustrate the benefits of SSSR. In this application, two functions use discrete optimization to explore different solutions and select the one leading to the minimal cost in terms of bitrate/quality and computational energy: coding-tree partitioning and intra-mode prediction. By applying SSSR to this use case, energy reductions from 20% to 70% are explored through Pareto in Rate-Energy space. Alexandre Mercat, Justine Bonnot, Maxime Pelcat, Wassim Hamidouche, Daniel Ménard |
DATE | 4 |
| 2017 | Cluster Adapted Signalling for Intra Prediction in HEVCabstractThe High Efficiency Video Coding (HEVC) standard defines 35 Intra Prediction Modes (IPM) to provide an efficient compression of intra coded blocks. Those IPMs are signalled to the decoder through the use of three compression tools: prediction, clustering and coding. In this paper we provide improvements to these three tools through: new labels for the prediction, new tests for the clustering and new coding schemes for the coding. The most significant improvement consists in the provision of a cluster-dependent code: adapting the coding scheme to the available information enables the average symbol cost to get within close margin of the entropy of the data. The system providing the best compression efficiency based on these improvements is then computed, enabling significant reduction in the average cost required to code the IPMs. The proposed method builds a new coding system with the same complexity as HEVC with 0.41% bit-rates savings in All Intra coding configuration. Kevin Reuze, Pierrick Philippe, Wassim Hamidouche, Olivier Déforges |
DCC | 3 |
| 2017 | Inter-block dependencies consideration for intra coding in H.264/AVC and HEVC standardsabstractRecent MPEG video compression standards are still block-based: blocks of pixels are sequentially coded using spatial or temporal prediction schemes. For each block, a vector of coding parameters has to be selected. In order to limit the complexity of this decision, independence between blocks is assumed, and coding parameters are locally optimized to maximize the coding efficiency. Few studies have investigated the benefits of inter-block dependencies consideration using Joint Rate-Distortion Optimization (JRDO), especially in Intra coding. To the best of our knowledge, maximum achievable gains of such approaches have never been exhibited. In this paper, we propose two JRDO models performing joint optimization of multiple blocks applied to intra prediction mode decision. The proposed models have been evaluated in both H.264/AVC and HEVC standards. These two models enables a bitrate saving with respect to the classical RDO model up to -3.10% and -2.31% in H.264/AVC and HEVC, respectively. Maxime Bichon, Julien Le Tanou, Michaël Ropert, Wassim Hamidouche, Luce Morin, Lu Zhang 0037 |
ICASSP | 4 |
| 2017 | Real-time and parallel SHVC hybrid codec AVC to HEVC decoderabstractScalable High efficiency Video Coding (SHVC) is the scalable extension of the latest video coding standard High Efficiency Video Coding (HEVC). One of the key novelties introduced by SHVC is that it enables hybrid codec scalability. This basically means that the video layers can be encoded with different video standards providing backward compatibility between codecs. In this paper, we propose a software parallel SHVC decoder in hybrid codec scalability configuration. The proposed design consists of an Advanced Video Coding (AVC) decoder for the Base Layer (BL) and a HEVC decoder for the Enhanced Layer (EL). In order to perform Inter Layer Prediction (ILP), a communication of decoding states and outputs is established between the two decoders. While the native frame based parallelism is still allowed within the two decoders, the proposed design also enables the use of frame based parallelism between the two decoders. The proposed software design enables a real time decoding of the HEVC EL at 2160p60 while the AVC base layer is decoded at 1080p60 for ×2 spatial scalability. Pierre-Loup Cabarat, Wassim Hamidouche, Olivier Déforges |
ICASSP | 2 |
| 2017 | Energy reduction opportunities in an HEVC real-time encoderabstractHigh Efficiency Video Coding (HEVC) is one of the latest released video standards and offers up to 40% bitrate savings when compared to the widespread H.264/AVC standard, at the cost of a substantial complexity growth. Constraining the complexity of HEVC encoding is a challenging task for embedded applications based on a software encoder. In the last few years, the Internet of Thingss (IoTs) has become a reality. Forecoming applications are likely to boost mobile video demand to an unprecedented level. In this context, designing energy-efficient HEVC real-time encoders is becoming a major challenge for software and hardware designers. In this paper, an analysis is conducted of the energy reduction opportunities offered by an HEVC encoder. The energy reduction search space is demonstrated, and the impact on energy consumption of encoding tools at various levels of granularity is measured. Alexandre Mercat, Florian Arrestier, Wassim Hamidouche, Maxime Pelcat, Daniel Ménard |
ICASSP | 3 |
| 2017 | Constrain the Docile CTUs: An In-Frame complexity allocator for HEVC Intra encodersabstractHigh Efficiency Video Coding (HEVC) is one of the latest released video standards and offers up to 40% bitrate savings when compared to the widespread H.264/AVC standard, at the cost of a substantial complexity growth. Constraining the complexity of HEVC encoding is a challenging task for embedded applications based on a software encoder. The most frequent approach to solve this problem is to optimise the coding tree structure to balance compression efficiency and computational complexity. In this context, we propose and assess a method to adequately allocate the computational complexity among coding units in a frame encoded in Intra mode. By studying an open-source real-time HEVC encoder, correlations are observed between Rate-Distortion (RD)-cost and encoding complexity that motivate a new complexity allocation technique. This technique, called “Constrain the Docile CTUs” (CDC), consists of allocating less computational complexity to units with low RD-costs and using RD-costs from preceding images as predictors for the current RD-costs. Experimental results demonstrate substantial gains, up to 36% of Bjøntegaard Delta Bit Rate (BD-BR), when using CDC method instead of other allocation methods. Alexandre Mercat, Florian Arrestier, Wassim Hamidouche, Maxime Pelcat, Daniel Ménard |
ICASSP | 3 |
| 2017 | A new perceptual assessment methodology for selective HEVC video encryptionabstractVideo data security is one of the most research topic in the recent years. It is widely used in the multimedia applications such as video-conferencing, Video on Demand and Pay-TV services. Although many video encryption methods and objective measurements have been employed, few real time schemes and no subjective studies have been proposed. In this paper we investigate a set of selective video encryption schemes by encrypting only a few parameters in HEVC video streams. Firstly, we carried out an in-depth subjective study of three proposed selective HEVC video encryption schemes. A panel of observers has participated in this test campaign in order to evaluate the degree of visibility of the encrypted videos at different bitrates. Experimental results are presented and analysed, showing therefore that two proposed selected encryption schemes allow a high perceptual security level by masking the whole details of the video content, while the third scheme achieves a high security level, with a content nearly unidentifiable. In addition, subjective scores can be used as ground truth for assessing selective video encryption methods, instead of classical objective signal-based metrics, which are not correlated with human judgment. Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges |
ICASSP | 2 |
| 2017 | An adaptive perceptual quantization method for HDR video codingabstractThis paper presents a new adaptive perceptual quantization method for the High Dynamic Range (HDR) content. This method considers the luminance distribution of the HDR image as well as the Minimum Detectable Contrast (MDC) thresholds to preserve the contrast information during quantization. Base on this method, we develop a mapping function for HDR video compression and apply it to a HEVC Main 10 Profile-based video coding chain. Our experiments show that the proposed mapping function can efficiently improve the quality of the reconstructed HDR video in both objective and subjective assessments. Yi Liu 0004, Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Giuseppe Valenzise, Emin Zerman |
ICIP | 3 |
| 2017 | Multi-output speckle reduction filter for ultrasound medical images based on multiplicative multiresolution decompositionabstractUltrasonographic examination, either as visual inspection or quantitative analysis, is less effective than other medical imaging systems due to speckle noise. The state-of-the-art speckle reduction methods often offers an effective speckle reduction but generally they suffer from oversmoothig, blurring effect and man-made/artificial appearance. In this paper, a new Multi-Output Filter based on a Multiplicative Multiresolution Decomposition (MOF-MMD) is proposed. This multiscale based method, particularly efficient in the case of multiplicative noise, enhances distinctively three outputs: edges, texture and the global image. The multi-output filter aims at offering an enhanced images according to the features desired by radiologists. The different structures, textures and edges are filtered according to the contour image obtained by morphological operators. Finally, we compare the MOF-MMD method with two state-of-the-art speckle reduction methods in terms of speckle reduction capacity and image quality improvement. The results show that the proposed method offers an effective speckle reduction with an improvement of the image quality without blurry and over-smoothing effect. Meriem Outtas, Lu Zhang 0037, Olivier Déforges, Amina Serir, Wassim Hamidouche |
ICIP | 5 |
| 2017 | Compression efficiency of the emerging video coding toolsabstractWith the drastic increasing of multimedia applications and video coarse consumption, video compression and content quality evaluation have become an exciting and challenging topic. Recently, a new coding tool has been developed under the Joint Exploration Model (JEM) software with the main goal to provide a high bit rate saving compared to the HEVC standard. In this paper we present a performance-based comparison between the JEM and HEVC reference software (HM) through an objective measurements and a subjective quality assessments. A set of video sequences, in two spatial resolutions High Definition (HD) and Ultra-High Definition (UHD), have been used in this study. These videos are encoded using both JEM and HM software at different bitrates. Results have shown that the JEM codec enables, subjectively, a quality enhancement up to 40% at similar low bit rates. Objectively, this quality improvement is ranging from 35% to 37% depending on the spatial resolution. However, at high bit rate, the HM reference software enables a high video quality and thus its becomes more difficult to perceive the quality enhancement is about brought by the JEM codec. In addition, some video contents are difficult to encode and, consequently, the JEM enables only slight perceived quality improvement especially at the highest considered bitrates and for 4K resolutions. Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Pierrick Philippe |
ICIP | 2 |
| 2017 | Emerging video coding performance: 4K quality monitoringabstractThe new coding tools, developed under the Joint Exploration Model (JEM) software, have been proposed with the main goal to explore their potential coding gain in the perspective to develop a new video coding standard. In this paper we present a performance-based comparison between the JEM and HEVC reference software (HM) through a set of subjective quality assessments. Different video sequences, encoded using both JEM and HM software at different bitrates, have been used in this experiment. Results have shown that the JEM codec enables, subjectively, a quality enhancement up to 40% at similar low bit rates. Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Pierrick Philippe |
QoMEX | 2 |
| 2017 | Smart search space reduction for approximate computing: A low energy HEVC encoder case study
Alexandre Mercat, Justine Bonnot, Maxime Pelcat, Karol Desnos, Wassim Hamidouche, Daniel Ménard |
J. Syst. Archit. | 5 |
| 2017 | Real-time selective video encryption based on the chaos system in scalable HEVC extension
Wassim Hamidouche, Mousa Farajallah, Naty Ould Sidaty, Safwan El Assad, Olivier Déforges |
Signal Process. Image Commun. | 1 |
| 2016 | Optimal Bitrate Allocation for High Dynamic Range and Wide Color Gamut Services Deployment Using SHVCabstractThe scalable video coding enables to compress video contents into a hierarchical layered representation, each layer depicts an enhanced version of the underlying layer. SHVC is the scalable extension of HEVC and enables spatial, SNR, color-gamut, codec and bitdepth scalability. It has been proved, in the MPEG investigations prior to the recent Call for Evidence, that SHVC can support SDR-to-HDR scalability by using the color gamut scalability, when SDR and HDR signals are placed in different color gamuts. This way, SHVC can be used to address future backward compatible issues in the HDR and WCG services deployment. In this paper, we exploit the impact of bitrate ratio over performance in scalable schemes to design an adaptive rate control algorithm suitable for such deployment, considering adjustable quality and bandwidth constraints. Our method dynamically adjusts the bitrate ratio between two layers during encoding in the most quality-related optimal way under specified constraints. The proposed method is tested on scalable combinations of HD/UHD, R.709/DCI-P3/R.2020 and SDR/HDR video contents, and reduces the average overhead introduced by SHVC compared to the single-layer HEVC encoding by 23%. Thibaud Biatek, Wassim Hamidouche, Jean-François Travers, Olivier Déforges |
DCC | 2 |
| 2016 | Adaptive rate control algorithm for SHVC: Application to HD/UHDabstractScalable video coding consists in compressing the video sequence into a layered bitstream where each layer refers to different spatial, temporal or quality representation of the video. Scalability enables compression gain compared to the simulcast encoding of layers thanks to inter-layer predictions. The scalable HEVC extension (SHVC) is the latest scalable technology promising up to 30% bitrate gains under the common test conditions, defined by JCT-VC. These conditions do not consider UHD and use fixed quantization step, which is not relevant in operational environment. In this paper, we propose an innovative adaptive rate control algorithm for SHVC. We consider HD as a base layer and UHD as an enhancement layer, with a constant global bitrate and a dynamic bitrate ratio adjustment between layers. The proposed algorithm is evaluated on a UHD data set where enables on average a BD-BR gain of 4.25% compared to a fixed-ratio encoding. Thibaud Biatek, Wassim Hamidouche, Jean-François Travers, Olivier Déforges |
ICASSP | 2 |
| 2016 | Generic statistical multiplexer with a parametrized bitrate allocation criteriaabstractIn this paper, we address the problem of the statistical multiplexing of video streams. Dynamic bitrate allocation is used to improve the overall video quality of a pool of channels. The balance is obtained by providing more bits to complex channels, while deprivations are applied to non-complex ones. In this study, the error minimization optimization of several compressed video is considered along with different metrics in order to exhibit a repartition key for bitrate sharing among all the channels. The goal of this approach is to introduce a reactivity parameter able to manage the bit transfer between channels. The validity of the parametric model is verified on two particular values, and compared to a static repartition solution. Médéric Blestel, Michaël Ropert, Wassim Hamidouche |
ICIP | 3 |
| 2016 | Pre-encoding based statistical-multiplexing for hybrid delivery of UHD services using SHVCabstractThe scalable video coding consists in encoding the video content into multiple representations, called layers, where each one refers to a specific version of the content. The scalable extension of the High Efficiency Video Coding standard SHVC is currently considered in ATSC and DVB to carry out layered and scalable programs which enables to target multiple equipments, and is also used to ensure backward compatibility with legacy receivers. In the case of hybrid delivery of HD and UHD services using SHVC, these services are encoded in two layers including base and enhancement layers, which are then broadcasted over separated channels. In this paper, a statistical multiplexing method is proposed for broadcasting of UHD services in this hybrid scenario. This innovative method considers both variable bitrate among programs and optimal SHVC layers coding, which was not considered in the existing approaches. The proposed method enables to reduce the overhead introduced by SHVC compared to the single-layer encoding by 3.3% in average while maintaining smooth quality variations among programs. Thibaud Biatek, Wassim Hamidouche, Jean-François Travers, Olivier Déforges |
PCS | 2 |
| 2016 | HDR video quality evaluation of HEVC and VP9 codecsabstractCurrent increasing effort in the television industry towards High Dynamic Range (HDR) imaging has raised the issue of the compression of HDR content. Offering a higher peak luminance and wider color gamut, HDR video introduces new challenges to the state-of-the-art video codecs such as High Efficiency Video Coding (HEVC) or VP9, which have been designed and optimized for the compression of Standard Dynamic Range (SDR) content. This study presents a performance comparison between HEVC and VP9 in the HDR context through both objective and subjective evaluations. The experimental objective results have shown that HEVC offers from 0.6% to 38.2% bit rate savings over VP9 depending on the objective metric which is used. The subjective study demonstrated that, on average, bit rate savings greater than 47.7% can be achieved by HEVC for the same perceived quality as VP9. Glenn Herrou, Wassim Hamidouche, Xavier Ducloux |
PCS | 2 |
| 2016 | Intra prediction modes signalling in HEVCabstractThe High Efficiency Video Coding (HEVC) standard defines 35 Intra Prediction Modes (IPM) to provide an efficient compression of intra coded blocks. To signal these IPMs to the decoder a list of 3 Most Probable Modes (MPM) is created based on the IPMs of the neighbour Intra coded blocks. These MPMs will be transmitted in the bit stream with a reduced number of bits. However, the signalling scheme used in HEVC is not optimal and can be further improved. In this paper a method is proposed to enhance the Intra mode signalling in HEVC. The IPM signaling is first modelled as a tree with forks and leaves representing tests and labels, respectively. The proposed solution introduces new decision tree process by adding new tests and new labels not considered in HEVC. This solution provides a systematic way to find the best signalling scheme for a given set of data. Experimental results show that the proposed solution enables to reduce the BD-Rate by 0.38% in all Intra coding configuration. Kevin Reuze, Pierrick Philippe, Olivier Déforges, Wassim Hamidouche |
PCS | 4 |
| 2016 | 4K Real-Time and Parallel Software Video Decoder for Multilayer HEVC ExtensionsabstractTwo High Efficiency Video Coding (HEVC) extensions, namely, the scalable HEVC (SHVC) extension and multiview HEVC (MV-HEVC) extension, have been finalized in July 2014 by the Moving Picture Experts Group and Video Coding Experts Group. These two extensions enable additional features not covered in the first version of the HEVC standard such as spatial, fidelity, bitdepth, and color gamut scalability, as well as stereoscopic and multiview representations. In this paper, we propose a software parallel decoder architecture for the HEVC standard and its multilayer extensions, including SHVC and MV-HEVC extensions. The decoder consists of multiple instances of the OpenHEVC decoder, one instance to decode each layer with a communication between dependent layers to perform inter-layer predictions. The proposed multilayer HEVC decoder is parallel friendly and supports both wavefront parallelism to simultaneously process adjacent rows of the frame and frame-based parallelism to decode a set of temporal and spatial frames in parallel. Moreover, the most time-consuming operation introduced in the SHVC extension, namely, the resampling of the inter-layer reference picture in spatial scalability, is optimized in single instruction multiple data for x86 platform. We assess the complexity of the multilayer HEVC decoder with respect to the simulcast configuration. The multilayer decoder decoding two SHVC layers introduces in average 40%-71% additional complexity compared with the single layer HEVC decoder. Moreover, the low level optimizations with a hybrid parallel processing solution enable a real-time decoding of 4Kp60 enhancement layer on a 6-core Intel i7 processor running at 3.4 GHz. Wassim Hamidouche, Mickaël Raulet, Olivier Déforges |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Selective video encryption using chaotic system in the SHVC extensionabstractIn this paper we investigate a selective video encryption in the scalable HEVC extension (SHVC). The SHVC extension encodes the video in several layers corresponding to different spatial and quality representations of the video. We propose a selective encryption solution using a chaotic-based encryption system. The proposed solution encrypts a set of sensitive parameters with a minimum complexity overhead, at constant bitrate and SHVC format compliant. Experimental results compare the performance of three encryption schemes: encrypt only the lowest layer, all layers, and only the highest layer. The first two schemes achieve a high security level with a drastic degradation in the decoded video, while the last scheme enables a perceptual video encryption by decreasing the quality of the highest layer below the quality of the clear layers. Wassim Hamidouche, Mousa Farajallah, Mickaël Raulet, Olivier Déforges, Safwan El Assad |
ICASSP | 1 |
| 2015 | ROI encryption for the HEVC coded video contentsabstractIn this paper we investigate privacy protection for the HEVC standard based on the tile concept. Tiles in HEVC enable the video to be split into independent rectangular regions. Two solutions are proposed to encrypt the tiles containing the Region Of Interest (ROI). The first solution performs encryption at the bitstream level by encrypting all HEVC syntax elements within the ROI tiles. The second solution enables a selective encryption of the ROI tiles under constant bitrate and format compliant requirements. To avoid temporal propagation of the encryption outside the ROI boundaries caused by inter prediction, the motion vectors of non ROI regions are restricted inside the non encrypted tiles in the reference frames. Simulation results show that the proposed solutions perform secure and adaptive encryption of ROI in the HEVC video. Moreover, the bitrate overhead caused by the MVs restriction window varies between 1%-2.5% depending on both the video content and the number of tiles within the frame. Mousa Farajallah, Wassim Hamidouche, Olivier Déforges, Safwan El Assad |
ICIP | 2 |
| 2014 | Multi-core software architecture for the scalable HEVC decoderabstractThe scalable high efficiency video coding (SHVC) standard aims to provide features of temporal, spatial and quality scalability. In this paper we investigate a pipeline and parallel software architecture for the SHVC decoder. The proposed architecture is based on the OpenHEVC software which implements the high efficiency video coding (HEVC) decoder. The architecture of the SHVC decoder enables two levels of parallelism. The first level decodes the base layer and the enhancement layers in parallel. The second level of parallelism performs the decoding of both the base layer and enhancement layers in parallel through the HEVC high level parallel processing solutions, including tile and wavefront. Up to the best of our knowledge, it is the first real time and parallel software implementation of the SHVC decoder. On an Intel Xeon processor running at 3.2 GHz, the SHVC decoder reaches the decoding of 1600p enhancement layer at 40 fps for x1.5 spatial scalability with using six concurent threads. Wassim Hamidouche, Mickaël Raulet, Olivier Déforges |
ICASSP | 1 |
| 2014 | Real time SHVC decoder: Implementation and complexity analysisabstractThe Scalable High efficiency Video Coding (SHVC) standard is developed to offer spatial and quality scalability with high coding efficiency. In this paper we investigate a complexity analysis of a real time and parallel SHVC decoder. We first provide details on the implementation of the SHVC decoder including its low level optimizations. Furthermore, we introduce parallelism tools integrated in the SHVC software for parallel decoding. These tools include frame-based parallelism to decode a set of temporal and spatial frames in parallel as well as wavefront parallelism to simultaneously process separated regions of a picture. We assessed through experimental results the complexity of the real time SHVC decoder in different coding configurations. The SHVC decoder with two layers introduces in average an additional complexity of 43 to 80% in respect to a simulcast configuration. The low level optimizations together with a hybrid parallelism solution enables a real time decoding of 1600p40 enhancement layer on an Intel i7 processor. Wassim Hamidouche, Mickaël Raulet, Olivier Déforges |
ICIP | 1 |
| 2014 | Unified real time software decoder for HEVC extensionsabstractThere are several High Efficiency Video Coding (HEVC) extensions providing new tools on the top of a common HEVC base layer. These HEVC extensions enable higher temporal, spatial or quality of the video, 3-D rendering and range extension. In addition to the conforming HEVC decoder, the end-user need to support the decoding of the HEVC extension to benefits from its tools. In this paper we propose an unified software decoder enabling to decode all HEVC extensions. This solution is based on the open source project OpenHEVC which implements a conforming HEVC decoder. The new tools defined in HEVC extension are implemented and integrated into the OpenHEVC decoder. We show the first end-to-end video streaming demonstration of the HEVC extensions with the OpenHEVC decoder and the GPAC player. The GPAC server streams one HEVC base layer and two enhancement layers. At the client side, GPAC player uses the OpenHEVC to decode the base layer for HD resolution and can also decode whether the first EL for 4K resolution or the second EL for 3-D rendering. Benoît Martin, Wassim Hamidouche, Jean Le Feuvre, Mickaël Raulet |
ICIP | 2 |
| 2014 | Parallel SHVC decoder: Implementation and analysisabstractThe new Scalable High efficiency Video Coding (SHVC) standard is based on a multi-loop coding structure which requires the total decoding of all intermediate layers. The decoding complexity becomes then a real issue, especially for a real time decoding of ultra high video resolutions. A parallel processing architecture is proposed to reduce both the decoding time and the latency of the SHVC decoder. The proposed solution combines the high level parallel processing solutions defined in the HEVC standard with an extension of the frame-based parallelism. The latter solution enables the decoding of several spatial and temporal SHVC frames in parallel to enhance both decoding frame rate and latency. The wavefront parallel processing solution is used for more coarse level of granularity. The proposed hybrid parallel processing approach achieves a near optimal speedup and provides a good trade-off between decoding time, latency and memory usage. On a 6 cores Xeon processor, the parallel SHVC decoder performs a real time decoding of 1600p60 video resolution. Wassim Hamidouche, Mickaël Raulet, Olivier Déforges |
ICME | 1 |
| 2013 | Optimal resource allocation for Medium Grain Scalable video transmission over MIMO channels
Wassim Hamidouche, Christian Olivier, Yannis Pousset, Clency Perrine |
J. Vis. Commun. Image Represent. | 1 |
| 2011 | A solution to efficient power allocation for H.264/SVC video transmission over a realistic MIMO channel using precoder designs
Wassim Hamidouche, Clency Perrine, Yannis Pousset, Christian Olivier |
J. Vis. Commun. Image Represent. | 1 |
| 2010 | Optimal solution for SVC-based video transmission over a realistic mimo channel using precoder designsabstractIn this paper we propose a novel scheme for SVC-based video transmission over MIMO channels using precoder solutions. On the one hand, H.264/SVC codec ensures spatial, temporal and quality scalabilities and provides an intrinsic hierarchy over transmitted bitstreams. On the other hand, precoder designs, which decouple a MIMO channel into parallel and independent SISOsub-channels, offer a high BER performance with a hierarchy over the sub-channels. The proposed scheme exploits the scalability of the H.264/SVC codec jointly with four precoder solutions providing an UEP without any surplus redundancy. Besides, some of these precoders such as QoS and E-dminallow a high flexibility on the power allocation across the sub-channels. This scalability is used to analytically draw the bandwidth allocation problem for H.264/SVC transmission over MIMO channels. The solution of this problem allows an optimal selection of each transmission bloc parameter approaching the optimal solution in term of Rate Distortion (RD) criterion. The simulation results show the performance of the proposed scheme over both statistical and realistic MIMO channels. Moreover, the accuracy of these precoder solutions against the Channel Estimation (CE) errors is investigated in different user mobility speeds. Wassim Hamidouche, Clency Perrine, Yannis Pousset, Christian Olivier |
ICASSP | 1 |
| 2009 | Impact of realistic MIMO physical layer on video transmission over mobile Ad Hoc networkabstractIn this paper we investigate the impact of a realistic physical layer on the H.264/AVC video transmission over Ad Hoc networks in urban environment. We propose a realistic Multiple Input Multiple Output (MIMO) physical layer which combines a determinist propagation model and a fine-grained model of wireless transmission errors. The determinist propagation model takes into account all the environmental characteristics (geometric and electric) and provides all the information of the multi-path channel (received power, complex impulse response). The wireless transmission errors model is based on a BER computation. The BER is calculated according to 802.11n standard to evaluate MIMO wireless links and, then, is compared to both SISO configuration and an existing wireless errors model using empirical propagation models. In the case of a SISO configuration, the BER is computed according to 802.11a standard. The simulation results show clearly a significant difference in term of QoS for the video transmission using realistic and empirical physical layer. In addition, the MIMO system, compared to a SISO one, improves the quality of links in the network and, thus, provides a better QoS for video transmission over Ad Hoc networks. Wassim Hamidouche, Rodolphe Vauzelle, Christian Olivier, Yannis Pousset, Clency Perrine |
PIMRC | 1 |