VLDB 2026 Research / reviewers in the wild / expert
Olivier Déforges
dblp:62/4354
· DBLP profile ↗
105ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0003-0750-0959ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 94 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 since 2021Systems, architecture and hardware · 6 · 1 since 2021Computer networks · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adversarial threats to vision transformers: evaluating robustness beyond CNNs
Ahmed Aldahdooh, Wassim Hamidouche, Olivier Déforges |
Neural Comput. Appl. | 3 |
| 2025 | Energy Backdoor Attack to Deep Neural NetworksabstractThe rise of deep learning (DL) has increased computing complexity and energy use, prompting the adoption of application specific integrated circuits (ASICs) for energy-efficient edge and mobile deployment. However, recent studies have demonstrated the vulnerability of these accelerators to energy attacks. Despite the development of various inference time energy attacks in prior research, backdoor energy attacks remain unexplored. In this paper, we design an innovative energy backdoor attack against deep neural networks (DNNs) operating on sparsity-based accelerators. Our attack is carried out in two distinct phases: backdoor injection and backdoor stealthiness. Experimental results using ResNet-18 and MobileNet-V2 models trained on CIFAR-10 and Tiny ImageNet datasets show the effectiveness of our proposed attack in increasing energy consumption on trigger samples while preserving the model’s performance for clean/regular inputs. This demonstrates the vulnerability of DNNs to energy backdoor attacks. The source code of our attack is available at: https://github.com/hbrachemi/energybackdoor. Hanene Brachemi Meftah, Wassim Hamidouche, Sid Ahmed Fezza, Olivier Déforges, Kassem Kallas |
ICASSP | 4 |
| 2025 | Improved Encoding for Overfitted Video CodecsabstractOverfitted neural video codecs offer a decoding complexity orders of magnitude smaller than their autoencoder counterparts. Yet, this low complexity comes at the cost of limited compression efficiency, in part due to their difficulty capturing accurate motion information. This paper proposes to guide motion information learning with an optical flow estimator. A joint rate-distortion optimization is also introduced to improve rate distribution across the different frames. These contributions maintain a low decoding complexity of 1300 multiplications per pixel while offering compression performance close to the conventional codec HEVC and outperforming other overfitted codecs. This work is made open-source at https://orange-opensource.github.io/Cool-Chic/. Thomas Leguay, Théo Ladune, Pierrick Philippe, Olivier Déforges |
ISCAS | 4 |
| 2025 | Siamese Network-Based Detection of Deepfake Impersonation Attacks with a Person of Interest ApproachabstractDeepfake technology presents critical cybersecurity challenges that have become more popular since easily accessible applications have become more widely available. The proliferation of fake portrait videos constitutes a serious risk to the legal system, society, and personal privacy. The publication of fraudulent explicit content starring celebrities, the circulation of fake political videos, and the use of faked impersonated videos as proof in court of law are all examples of the effects of deepfakes in the real world. In reaction to this growing threat, we propose a simple yet efficient Person of Interest (PoI) Siamese network-based model to detect deepfake synthetic content in portrait images, providing a preventative measure against the growing danger of deepfakes. On one side and unlike traditional neural networks, which process inputs independently, our approach leverages a Siamese network that processes two inputs simultaneously using identical sub-networks with shared weights and parameters. This twin structure is particularly effective for our adopted PoI methodology, where one input is a reference image of a specific individual, and the other is an image that needs to be verified as either real or fake. By ensuring that both the reference and the suspect image are processed in the same way, the network can accurately learn and detect subtle differences, enabling it to determine whether the second image is a genuine representation of the individual or a deepfake. This makes our method particularly relevant for targeted forensic investigations or security applications where an individual’s media integrity is paramount. On the other side, our proposed method does not require any additional complex biological feature extraction. Despite this simplification, our method achieves comparable accuracy to more complex models that rely on biological feature extraction. This efficiency makes our approach practical for implementation on resource-constrained devices, such as mobile phones and Internet of Things (IoT) systems. Khouloud Samrouth, Pia El Housseini, Olivier Déforges |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Cool-chic video: Learned video coding with 800 parametersabstractWe propose a lightweight learned video codec with 900 multiplications per decoded pixel and 800 parameters overall. To the best of our knowledge, this is one of the neural video codecs with the lowest decoding complexity. It is built upon the overfitted image codec Cool-chic and supplements it with an inter coding module to leverage the video’s temporal redundancies. The proposed model is able to compress videos using both low-delay and random access configurations and achieves rate-distortion close to AVC while outperforming other overfitted codecs such as FFNeRV. The system is made open-source: orange-opensource.github.io/Cool-Chic. Thomas Leguay, Théo Ladune, Pierrick Philippe, Olivier Déforges |
DCC | 4 |
| 2024 | Res-NeRV: Residual Blocks For A Practical Implicit Neural Video DecoderabstractThis paper proposes the integration of residual blocks into neural representation for videos (NeRV)-based architectures with the aim of enhancing the reconstruction of detailed patterns and high-level features. Additionally, a coding pipeline is introduced, placing the implicit neural decoder in a real-life video streaming framework. Indeed, DeepCABAC is employed for model compression, applying a quantization scheme followed by the context-adaptive binary arithmetic coding (CABAC) entropy coding algorithm, ultimately leading to bitstream generation. Our method outperforms NeRV, as well as x264 and x265, achieving BD-rate gains against NeRV: $-12.06 \%$ using PSNR and $-14.25 \%$ using MS-SSIM. Furthermore, it exhibits superior subjective quality compared to NeRV, attributed to enhanced high-level feature reconstruction. This observed behavior encourages the application of our method to other NeRV-based models, such as E-NeRV. Marwa Tarchouli, Thomas Guionnet, Marc Rivière, Wassim Hamidouche, Meriem Outtas, Olivier Déforges |
ICIP | 6 |
| 2024 | Hierarchical Learning and Dummy Triplet Loss for Efficient Deepfake DetectionabstractThe advancement of generative models has made it easier to create highly realistic Deepfake videos. This accessibility has led to a surge in research on Deepfake detection to mitigate potential misuse. Typically, Deepfake detection models utilize binary backbones, even though the training dataset contains additional exploitable information, such as the Deepfake generation method employed for each video. However, recent findings suggest that inferring a binary class from a multi-class backbone yields superior performance compared to directly employing a binary backbone. Building upon this research, our article introduces two novel methods to infer a binary class from a multi-class backbone. The first method, named root dummies , leverages the dummy triplet loss, which employs fixed vectors (i.e., dummies) instead of mined positives and negatives in the triplet loss. By training the multi-class backbone with these dummies, we can easily infer a binary class during testing by adjusting the number of dummies (from six during training to two during inference). Through this approach, we achieve an accuracy improvement of 0.23% compared to the existing inference method, without requiring additional training. The second proposed method is transfer learning. It involves training a classifier, such as a support vector machine, to predict binary classes based on the image embeddings generated by the multi-class backbone. Although this method necessitates additional training, it further enhances the model’s performance, resulting in an accuracy increase of 1.79%. In summary, our proposed methods improve the accuracy of Deepfake detection by simply modifying the number of classes during training, making them suitable for integration into a variety of existing Deepfake training pipelines. Additionally, to foster reproducible research, we have made the source code of our solution publicly available at https://github.com/beuve/DmyT . Nicolas Beuve, Wassim Hamidouche, Olivier Déforges |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Energy Efficient VVC Decoding on Mobile PlatformabstractRecently, global demand for high-resolution videos and new multimedia applications have created the need for a new video coding standard. Hence, in July 2020 the Versatile Video Coding (VVC) standard was released providing up to 40% bit-rate saving for the same video quality compared to its predecessor High Efficiency Video Coding (HEVC). However, this bit-rate saving comes at the cost of high computational complexity, particularly for live applications and on resource-constraint embedded devices. This paper presents an power-efficient VVC decoder implementation designed for low-resource platforms. This latter exploits optimization techniques such as data level parallelism using Single Instruction Multiple Data (SIMD) instructions and functional level parallelism using frame, tile and slice-based parallelisms. The results showed that the OpenVVC decoder achieve real-time decoding of Full High Definition (FHD) resolution at 30 fps targeting a platform with 8 cores with a maximum frequency of 2.2 Ghz and High Definition (HD) real-time decoding at 30 fps for platforms using 4 cores with a maximum frequency of 1.8 Ghz. In terms of average consumed power, OpenVVC showed around 5.6 watts and 1.7 watts for the 8 and 4 cores platforms, respectively. In addition, it comes with the best trade-off between what is achievable in real-time and the power consumed in comparison to the state-of-the-art implementation. Ibrahim Farhat, Pierre-Loup Cabarat, Daniel Ménard, Wassim Hamidouche, Olivier Déforges |
MMSP | 5 |
| 2023 | Low-Complexity Overfitted Neural Image CodecabstractWe propose a neural image codec at reduced complexity which overfits the decoder parameters to each input image. While autoencoders perform up to a million multiplications per decoded pixel, the proposed approach only requires 2300 multiplications per pixel. Albeit low-complexity, the method rivals autoencoder performance and surpasses HEVC performance under various coding conditions. Additional lightweight modules and an improved training process provide a 14% rate reduction with respect to previous overfitted codecs, while offering a similar complexity. This work is made open-source at http://orange-opensource.github.io/Cool-Chic/. Thomas Leguay, Théo Ladune, Pierrick Philippe, Gordon Clare, Félix Henry, Olivier Déforges |
MMSP | 6 |
| 2023 | Denoised CT Images Quality Assessment Through COVID-19 Pneumonia Detection TaskabstractMedical images largely contribute to the diagnosis of lung diseases, especially pneumonia, an inflammation of lungs tissue. Since the emergence of COVID-19 in late 2019, medical imaging systems, notably computed tomography (CT) scans, have considerably helped in its diagnosis as well as revealing its infection severity. Serving as such an important role in clinical practice, the quality of medical images is therefore crucial for an accurate diagnosis. Denoising techniques, as a common image processing method, are being more and more used in medical imaging. However, how image denoising technique influences medical images' quality in terms of diagnostic performance still remains to be answered. In this paper, a primary study was carried out thanks to a detection task-based image quality assessment experiment, where we explored the performance of COVID-19 classifiers on both original and denoised chest CT scans. Two different denoising methods, i.e., anisotropic diffusion (AD) and total variation (TV) filters, were used. Results showed that the TV denoised model performed better than both baseline and AD denoised model, despite its less favorable mathematical image quality metrics. Lumi Xia, Houda Jebbari, Olivier Déforges, Lu Zhang 0037, Lucie Lévêque, Meriem Outtas |
QoMEX | 3 |
| 2023 | Revisiting model's uncertainty and confidences for adversarial example detection
Ahmed Aldahdooh, Wassim Hamidouche, Olivier Déforges |
Appl. Intell. | 3 |
| 2023 | Predictive Uncertainty Estimation for Camouflaged Object DetectionabstractUncertainty is inherent in machine learning methods, especially those for camouflaged object detection aiming to finely segment the objects concealed in background. The strong enquote center bias of the training dataset leads to models of poor generalization ability as the models learn to find camouflaged objects around image center, which we define as enquote model bias. Further, due to the similar appearance of camouflaged object and its surroundings, it is difficult to label the accurate scope of the camouflaged object, especially along object boundaries, which we term as enquote data bias. To effectively model the two types of biases, we resort to uncertainty estimation and introduce predictive uncertainty estimation technique, which is the sum of model uncertainty and data uncertainty, to estimate the two types of biases simultaneously. Specifically, we present a predictive uncertainty estimation network (PUENet) that consists of a Bayesian conditional variational auto-encoder (BCVAE) to achieve predictive uncertainty estimation, and a predictive uncertainty approximation (PUA) module to avoid the expensive sampling process at test-time. Experimental results show that our PUENet achieves both highly accurate prediction, and reliable uncertainty estimation representing the biases within both model parameters and the datasets. Yi Zhang 0076, Jing Zhang 0052, Wassim Hamidouche, Olivier Déforges |
IEEE Trans. Image Process. | 4 |
| 2023 | PAV-SOD: A New Task towards Panoramic Audiovisual Saliency DetectionabstractObject-level audiovisual saliency detection in 360° panoramic real-life dynamic scenes is important for exploring and modeling human perception in immersive environments, also for aiding the development of virtual, augmented, and mixed reality applications in fields such as education, social network, entertainment, and training. To this end, we propose a new task, p anoramic a udio v isual s alient o bject d etection, ( PAV-SOD 1 ), which aims to segment the objects grasping most of the human attention in 360° panoramic videos reflecting real-life daily scenes. To support the task, we collect PAVS10K , the first p anoramic video dataset for a udio v isual s alient object detection, which consists of 67 4K-resolution equirectangular videos with per-video labels including hierarchical scene categories and associated attributes depicting specific challenges for conducting PAV-SOD , and 10,465 uniformly sampled video frames with manually annotated object-level and instance-level pixel-wise masks. The coarse-to-fine annotations enable multi-perspective analysis regarding PAV-SOD modeling. We further systematically benchmark 13 state-of-the-art salient object detection (SOD)/video object segmentation (VOS) methods based on our PAVS10K . Besides, we propose a new baseline network, which takes advantage of both visual and audio cues of 360° video frames by using a new conditional variational auto-encoder (CVAE). Our C VAE-based a udio v isual net work, namely, CAV-Net , consists of a spatial-temporal visual segmentation network, a convolutional audio-encoding network, and audiovisual distribution estimation modules. As a result, our CAV-Net outperforms all competing models and is able to estimate the aleatoric uncertainties within PAVS10K . With extensive experimental results, we gain several findings about PAV-SOD challenges and insights towards PAV-SOD model interpretability. We hope that our work could serve as a starting point for advancing SOD towards immersive media. Yi Zhang 0076, Fang-Yi Chao, Wassim Hamidouche, Olivier Déforges |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Channel-Spatial Mutual Attention Network for 360° Salient Object DetectionabstractIn this work, we conduct 360° panoramic salient object detection by taking advantage of both the global and local visual cues of 360° images, with a novel channel-spatial mutual attention network (CSMA-Net). The key component of the CSMA-Net is the proposed CSMA module, which cascades channel-/spatial-weighting-based mutual attentions. The objective of our CSMA module is to refine and fuse the bottleneck features from two separate encoders with different planar representations of 360° panorama as inputs, i.e., equirectangular image and cube map. Our CSMA-Net outperforms 10 state-of-the-art segmentation methods based on the proposed 360° SOD benchmark where multiple fine-tuning and testing strategies are applied to the widely-used 360° datasets. Extensive experimental results illustrate the effectiveness and robustness of the proposed CSMA-Net1. Yi Zhang 0076, Wassim Hamidouche, Olivier Déforges |
ICPR | 3 |
| 2022 | Efficient HW Design of Adaptive Loop Filter for 4k ASIC VVC EncoderabstractVersatile Video Coding (VVC) is the next-generation video coding standard released in July 2020. VVC introduces new coding tools enhancing the coding efficiency compared to its predecessor, High Efficiency Video Coding (HEVC). These new tools significantly impact the VVC software and hardware implementations with a complexity estimated to two times and eight times the HEVC decoder and encoder complexity, respectively. In particular, the Adaptive Loop Filter (ALF), adopted in VVC as an in-loop filter, increases both the run time complexity and memory usage. These concerns need to be carefully addressed regarding the design of a VVC hardware encoder. In this paper, we present an efficient hardware implementation of the ALF tool with its decision process in the context of a professional VVC encoder. The proposed solution can reach a real-time encoding of 4K resolution videos with 4:2:2 chroma sub-sampling at 60 frames per second targeting ASIC platforms with 28-nm technology. Ibrahim Farhat, Wassim Hamidouche, Adrien Grill, Daniel Ménard, Olivier Déforges |
PCS | 5 |
| 2022 | Visual Attention-Aware High Dynamic Range Quantization for HEVC Video CodingabstractEmerging HDR videos enable the recording of adequate luminance information and representing realistic scenes to the audience. Due to the high precision of the HDR data recorded in the floating-point format, a quantization process is required to convert HDR data to integer data for compatibility with current transmission and display systems. In this study, a novel attention-aware quantization method is presented that attempts to preserve the contrast details in the region of interest of the human visual system. This method was applied in the context of HDR video coding. The proposed coding solution was compared with the current anchor solution in terms of the quality of the reconstructed video. Experimental results show that the proposed solution is able to improve the visual quality of encoded video with respect to the anchor solution. Additionally, the proposed solution achieves a bit-rate gain over the anchor with reference to the objective evaluation results. Yi Liu 0004, Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Cheolkon Jung |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Learning Synergistic Attention for Light Field Salient Object Detection
Yi Zhang 0076, Geng Chen 0001, Yong Xia 0001, Olivier Déforges, Wassim Hamidouche, Lu Zhang 0037 |
BMVC | 6 |
| 2021 | Multitask Learning for VVC Quality Enhancement and Super-ResolutionabstractThe latest video coding standard, called versatile video coding (VVC), includes several novel and refined coding tools at different levels of the coding chain. These tools bring significant coding gains with respect to the previous standard, high efficiency video coding (HEVC). However, the encoder may still introduce visible coding artifacts, mainly caused by coding decisions applied to adjust the bitrate to the available bandwidth. Hence, pre and post-processing techniques are generally added to the coding pipeline to improve the quality of the decoded video. These methods have recently shown outstanding results compared to traditional approaches, thanks to the recent advances in deep learning. Generally, multiple neural networks are trained independently to perform different tasks, thus omitting to benefit from the redundancy that exists between the models. In this paper, we investigate a learning-based solution as a post-processing step to enhance the decoded VVC video quality. Our method relies on multitask learning to perform both quality enhancement and super-resolution using a single shared network optimized for multiple degradation levels. The proposed solution enables a good performance in both mitigating coding artifacts and super-resolution with fewer network parameters compared to traditional specialized architectures. Charles Bonnineau, Wassim Hamidouche, Jean-François Travers, Naty Ould Sidaty, Olivier Déforges |
PCS | 5 |
| 2021 | CAESR: Conditional Autoencoder and Super-Resolution for Learned Spatial ScalabilityabstractIn this paper, we present CAESR, an hybrid learning-based coding approach for spatial scalability based on the versatile video coding (VVC) standard. Our framework considers a low-resolution signal encoded with VVC intra-mode as a base-layer (BL), and a deep conditional autoencoder with hyperprior (AE-HP) as an enhancement-layer (EL) model. The EL encoder takes as inputs both the upscaled BL reconstruction and the original image. Our approach relies on conditional coding that learns the optimal mixture of the source and the upscaled BL image, enabling better performance than residual coding. On the decoder side, a super-resolution (SR) module is used to recover high-resolution details and invert the conditional coding process. Experimental results have shown that our solution is competitive with the VVC full-resolution intra coding while being scalable. Charles Bonnineau, Wassim Hamidouche, Jean-François Travers, Naty Ould Sidaty, Jean-Yves Aubié, Olivier Déforges |
VCIP | 6 |
| 2021 | Quality assessment of DIBR-synthesized views: An overview
Shishun Tian, Lu Zhang 0037, Wenbin Zou, Xia Li 0006, Ting Su 0004, Luce Morin, Olivier Déforges |
Neurocomputing | 7 |
| 2021 | Light Field Image Coding Using VVC Standard and View Synthesis Based on Dual Discriminator GAN
Nader Bakir, Wassim Hamidouche, Sid Ahmed Fezza, Khouloud Samrouth, Olivier Déforges |
IEEE Trans. Multim. | 5 |
| 2021 | A Multi-FoV Viewport-Based Visual Saliency Model Using Adaptive Weighting Losses for 360$^\circ$ Imagesabstract360$^\circ$media allows observers to explore the scene in all directions. The consequence is that the human visual attention is guided by not only the perceived area in the viewport but also the overall content in 360$^\circ$. In this paper, we propose a method to estimate the 360$^\circ$saliency map which extracts salient features from the entire 360$^\circ$image in each viewport in three different Field of Views (FoVs). Our model is first pretrained with a large-scale 2D image dataset to enable the interpretation of semantic contents, then fine-tuned with a relative small 360$^\circ$image dataset. A novel weighting loss function attached with stretch weighted maps is introduced to adaptively weight the losses of three evaluation metrics and attenuate the impact of stretched regions in equirectangular projection during training process. Experimental results demonstrate that our model achieves better performance with the integration of three FoVs and its diverse viewport images. Results also show that the adaptive weighting losses and stretch weighted maps effectively enhance the evaluation scores compared to the fixed weighting losses solutions. Comparing to other state of the art models, our method surpasses them on three different datasets and ranks the top using 5 performance evaluation metrics on the Salient360! benchmark set. The code is available athttps://github.com/FannyChao/MV-SalGAN360. Fang-Yi Chao, Lu Zhang 0037, Wassim Hamidouche, Olivier Déforges |
IEEE Trans. Multim. | 4 |
| 2020 | Versatile Video Coding and Super-Resolution for Efficient Delivery of 8k Video with 4k Backward-CompatibilityabstractIn this paper, we propose, through an objective study, to compare and evaluate the performance of different coding approaches allowing the delivery of an 8K video signal with 4K backward-compatibility on broadcast networks. Presented approaches include simulcast of 8K and 4K single-layer signals encoded using High-Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC) standards, spatial scalability using SHVC with 4K base layer (BL) and 8K enhancement-layer (EL), and super-resolution applied on 4K VVC signal after decoding to reach 8K resolution. For up-scaling, we selected the deep-learning-based super-resolution method called Super-Resolution with Feedback Network (SRFBN) and the Lanczos interpolation filter. We show that the deep-learning-based approach achieves visual quality gain over simulcast, especially on bit-rates lower than 30Mb/s with average gain of 0.77dB, 0.015, and 7.97 for PSNR, SSIM, and VMAF, respectively and outperforms the Lanczos filter in average by 29% of BD-rate savings. Charles Bonnineau, Wassim Hamidouche, Jean-François Travers, Olivier Déforges |
ICASSP | 4 |
| 2020 | Lightweight Hardware Implementation of VVC Transform Block for ASIC DecoderabstractVersatile Video Coding (VVC) is the next generation video coding standard expected by the end of 2020. Compared to its predecessor, VVC introduces new coding tools to make compression more efficient at the expense of higher computational complexity. This rises a need to design an efficient and optimised implementation especially for embedded platforms with limited memory and logic resources. One of the newly introduced tools in VVC is the Multiple Transform Selection (MTS). This latter involves three Discrete Cosine Transform (DCT)/Discrete Sine Transform (DST) types with larger and rectangular transform blocks. In this paper, an efficient hardware implementation of all DCT/DST transform types and sizes is proposed. The proposed design uses 32 multipliers in a pipelined architecture which targets an ASIC platform. It consists in a multi-standard architecture that supports the transform block of recent MPEG standards including AVC, HEVC and VVC. The architecture is optimized and removes unnecessary complexities found in other proposed architectures by using regular multipliers instead of multiple constant multipliers. The synthesized results show that the proposed method which sustain a constant throughput of two pixels/cycle and constant latency for all block sizes can reach an operational frequency of 600 Mhz enabling to decode in real-time 4K videos at 48 fps. Ibrahim Farhat, Wassim Hamidouche, Adrien Grill, Daniel Ménard, Olivier Déforges |
ICASSP | 5 |
| 2020 | Binary Probability Model for Learning Based Image CompressionabstractIn this paper, we propose to enhance learned image compression systems with a richer probability model for the latent variables. Previous works model the latents with a Gaussian or a Laplace distribution. Inspired by binary arithmetic coding, we propose to signal the latents with three binary values and one integer, with different probability models.A relaxation method is designed to perform gradient-based training. The richer probability model results in a better entropy coding leading to lower rate. Experiments under the Challenge on Learned Image Compression (CLIC) test conditions demonstrate that this method achieves 18 % rate saving compared to Gaussian or Laplace models. Théo Ladune, Pierrick Philippe, Wassim Hamidouche, Lu Zhang 0037, Olivier Déforges |
ICASSP | 5 |
| 2020 | A Fixation-Based 360° Benchmark Dataset For Salient Object DetectionabstractFixation prediction (FP) in panoramic contents has been widely investigated along with the booming trend of virtual reality (VR) applications. However, another issue within the field of visual saliency, salient object detection (SOD), has been seldom explored in 360° or omnidirectional) images due to the lack of datasets representative of real scenes with pixel-level annotations. Toward this end, we collect 107 equirectangular panoramas with challenging scenes and multiple object classes. Based on the consistency between FP and explicit saliency judgements, we further manually annotate 1,165 salient objects over the collected images with precise masks under the guidance of real human eye fixation maps. Six state-of-the-art SOD models are then benchmarked on the proposed fixation-based 360° image dataset (F-360iSOD), by applying a multiple cubic projection-based fine-tuning method. Experimental results show a limitation of the current methods when used for SOD in panoramic images, which indicates the proposed dataset is challenging. Key issues for 360° SOD is also discussed. The proposed dataset is available at https://github.com/PanoAsh/F-360iSOD. Yi Zhang 0076, Lu Zhang 0037, Wassim Hamidouche, Olivier Déforges |
ICIP | 4 |
| 2020 | Light Field Image Coding Using Dual Discriminator Generative Adversarial Network And VVC Temporal ScalabilityabstractLight field technology represents a viable path for providing a high-quality VR content. However, such an imaging system generates a high amount of data leading to an urgent need for LF image compression solution. In this paper, we propose an efficient LF image coding scheme based on view synthesis. Instead of transmitting all the LF views, only some of them are coded and transmitted, while the remaining views are dropped. The transmitted views are coded using Versatile Video Coding (VVC) and used as reference views to synthesize the missing views at decoder side. The dropped views are generated using the efficient dual discriminator GAN model. The selection of reference/dropped views is performed using a rate distortion optimization based on the VVC temporal scalability. Experimental results show that the proposed method provides high coding performance and overcomes the state-of-the-art LF image compression solutions. Nader Bakir, Wassim Hamidouche, Sid Ahmed Fezza, Khouloud Samrouth, Olivier Déforges |
ICME | 5 |
| 2020 | Detection of Adversarial Examples in Deep Neural Networks with Natural Scene StatisticsabstractRecent studies have demonstrated that the deep neural networks (DNNs) are vulnerable to carefully-crafted perturbations added to a legitimate input image. Such perturbed images are called adversarial examples (AEs) and can cause DNNs to misclassify. Consequently, it is of paramount importance to develop detection methods of AEs, thus allowing to reject them. In this paper, we propose to characterize the AEs through the use of natural scene statistics (NSS). We demonstrate that these statistical properties are altered by the presence of adversarial perturbations. Based on this finding, we propose three different methods that exploit these scene statistics to determine if an input is adversarial or not. The proposed detection methods have been evaluated against four prominent adversarial attacks and on three standards datasets. The experimental results have shown that the proposed methods achieve a high detection accuracy while providing a low false positive rate. Anouar Kherchouche, Sid Ahmed Fezza, Wassim Hamidouche, Olivier Déforges |
IJCNN | 4 |
| 2020 | Extending 2D Saliency Models for Head Movement Prediction in 360-Degree Images using CNN-Based FusionabstractSaliency prediction can be of great benefit for 360-degree image/video applications, including compression, streaming, rendering and viewpoint guidance. It is therefore quite natural to adapt the 2D saliency prediction methods for 360-degree images. To achieve this, it is necessary to project the 360-degree image to 2D plane. However, the existing projection techniques introduce different distortions, which provides poor results and makes inefficient the direct application of 2D saliency prediction models to 360-degree content. Consequently, in this paper, we propose a new framework for effectively applying any 2D saliency prediction method to 360-degree images. The proposed framework particularly includes a novel convolutional neural network based fusion approach that provides more accurate saliency prediction while avoiding the introduction of distortions. The proposed framework has been evaluated with five 2D saliency prediction methods, and the experimental results showed the superiority of our approach compared to the use of weighted sum or pixel-wise maximum fusion methods. Ibrahim Djemai, Sid Ahmed Fezza, Wassim Hamidouche, Olivier Déforges |
ISCAS | 4 |
| 2020 | Natural Scene Statistics for Detecting Adversarial Examples in Deep Neural NetworksabstractThe deep neural networks (DNNs) have been adopted in a wide spectrum of applications. However, it has been demonstrated that their are vulnerable to adversarial examples (AEs): carefully-crafted perturbations added to a clean input image. These AEs fool the DNNs which classify them incorrectly. Therefore, it is imperative to develop a detection method of AEs allowing the defense of DNNs. In this paper, we propose to characterize the adversarial perturbations through the use of natural scene statistics. We demonstrate that these statistical properties are altered by the presence of adversarial perturbations. Based on this finding, we design a classifier that exploits these scene statistics to determine if an input is adversarial or not. The proposed method has been evaluated against four prominent adversarial attacks and on three standards datasets. The experimental results have shown that the proposed detection method achieves a high detection accuracy, even against strong attacks, while providing a low false positive rate. Anouar Kherchouche, Sid Ahmed Fezza, Wassim Hamidouche, Olivier Déforges |
MMSP | 4 |
| 2020 | Optical Flow and Mode Selection for Learning-based Video CodingabstractThis paper introduces a new method for inter-frame coding based on two complementary autoencoders: MOFNet and CodecNet. MOFNet aims at computing and conveying the Optical Flow and a pixel-wise coding Mode selection. The optical flow is used to perform a prediction of the frame to code. The coding mode selection enables competition between direct copy of the prediction or transmission through CodecNet.The proposed coding scheme is assessed under the Challenge on Learned Image Compression 2020 (CLIC20) P-frame coding conditions, where it is shown to perform on par with the state-of-the-art video codec ITU/MPEG HEVC. Moreover, the possibility of copying the prediction enables to learn the optical flow in an end-to-end fashion i.e. without relying on pre-training and/or a dedicated loss term. Théo Ladune, Pierrick Philippe, Wassim Hamidouche, Lu Zhang 0037, Olivier Déforges |
MMSP | 5 |
| 2020 | Towards Audio-Visual Saliency Prediction for Omnidirectional Video with Spatial AudioabstractOmnidirectional videos (ODVs) with spatial audio enable viewers to perceive 360° directions of audio and visual signals during the consumption of ODVs with head-mounted displays (HMDs). By predicting salient audio-visual regions, ODV systems can be optimized to provide an immersive sensation of audio-visual stimuli with high-quality. Despite the intense recent effort for ODV saliency prediction, the current literature still does not consider the impact of auditory information in ODVs. In this work, we propose an audio-visual saliency (AVS360) model that incorporates 360° spatial-temporal visual representation and spatial auditory information in ODVs. The proposed AVS360 model is composed of two 3D residual networks (ResNets) to encode visual and audio cues. The first one is embedded with a spherical representation technique to extract 360° visual features, and the second one extracts the features of audio using the log mel-spectrogram. We emphasize sound source locations by integrating audio energy map (AEM) generated from spatial audio description (i.e., ambisonics) and equator viewing behavior with equator center bias (ECB). The audio and visual features are combined and fused with AEM and ECB via attention mechanism. Our experimental results show that the AVS360 model has significant superiority over five state-of-the-art saliency models. To the best of our knowledge, it is the first w ork that develops the audio-visual saliency model in ODVs. The code will be publicly available to foster future research on audio-visual saliency in ODVs. Fang-Yi Chao, Cagri Ozcinar, Lu Zhang 0037, Wassim Hamidouche, Olivier Déforges, Aljoscha Smolic |
VCIP | 5 |
| 2020 | Forward-Inverse 2D Hardware Implementation of Approximate Transform Core for the VVC StandardabstractThe future video coding standard named Versatile Video Coding (VVC) is expected by the end of 2020. VVC will enable better coding efficiency than the current High Efficiency Video Coding (HEVC) standard. This coding gain is brought by several coding tools. The Multiple Transform Selection (MTS) is one of the key coding tools that have been introduced in VVC. The MTS concept relies on three transform types including Discrete Cosine Transform (DCT)-II, Discrete Sine Transform (DST)-VII and DCT-VIII. Unlike the DCT-II that has fast computing algorithms, the DST-VII and DCT-VIII rely on more complex matrix multiplication. In this paper an approximation approach is proposed to reduce the computational cost of the DST-VII and DCT-VIII. The approximation consists in applying adjustment stages, based on sparse block-band matrices, to a variant of DCT-II family mainly DCT-II and its inverse. Genetic algorithm is used to derive the optimal coefficients of the adjustment matrices. Moreover, an efficient hardware implementation of the forward and inverse approximate transform module is proposed. The architecture design includes a pipelined and reconfigurable forward-inverse DCT-II core transform as it is the main core for DST-VII and DCT-VIII computations. The proposed 32-point 1D architecture including low cost adjustment stages allows the processing of a video in 2K and 4K resolutions at 1095 and 273 frames per second, respectively. A unified 2D implementation of forward-inverse DCT-II, approximate DST-VII and DCT-VIII is also presented. The synthesis results show that the design is able to sustain a video in 2K and 4K resolutions at 386 and 96 frames per second, respectively, while using only 12% of Alms, 22% of registers and 30% of DSP blocks of the Arria10 SoC platform. Ahmed Kammoun, Wassim Hamidouche, Pierrick Philippe, Olivier Déforges, Fatma Belghith, Nouri Masmoudi, Jean-François Nezan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | RDO-Based Light Field Image Coding Using Convolutional Neural Networks and Linear ApproximationabstractThe increasing penetration of acquisition and display devices for Light Field (LF) content in the consumer market leads to the high proliferation of this new immersive media. This growing interest to LF images thus urgently raises the question of their compression. In this paper, we propose a convolutional neural networks (CNN)-based LF image coding scheme including both Rate Distortion Optimization (RDO) and post-processing steps. First, at the encoder side, the views are rearranged in sparse and dropped set of views. The former are compressed with a standard encoder and transmitted, while the dropped views are either linearly approximated or synthesized by a CNN using the encoded views as input. This choice is made on the basis of the proposed RDO process. At the decoder side, once the dropped views are either linearly approximated or synthesized by a CNN block, a post-processing step is performed to further enhance the quality of the reconstructed views. This post-processing block is based on superpixel to pixel-matching. Experimental results show that the proposed scheme provides views with high visual quality and overcomes the state-of-the-art LF image compression solutions by -30% in terms of BD-BR and 0.62 dB in BD-PSNR. Nader Bakir, Wassim Hamidouche, Olivier Déforges, Khouloud Samrouth, Sid Ahmed Fezza |
DCC | 3 |
| 2019 | Dynamic Lists for Efficient Coding of Intra Prediction Modes in the Future Video Coding StandardabstractThe next generation MPEG video coding standard is under development by the Joint Video Coding Experts Team (JVET). This new standard, called Versatile Video Coding (VVC), is expected by the end of 2020 and will offer better coding efficiency than its predecessor High Efficiency Video Coding (HEVC) standard. This coding gain is enabled by new coding tools such as more flexible block partitioning, more accurate Intra/Inter predictions, multiple transforms and adaptive in-loop filtering. In this paper we focus on the coding of the Intra Prediction Modes (IPM) that have been increased from 35 modes in HEVC to 67 modes in VVC. We propose a solution based on genetic algorithms to build an ordered list for the coding of IPM in the Joint Exploration Model (JEM) codec. We first give the theoretical upper bound performance in terms of required bits per IPM to encode the IPM using the available contextual information. The new ordering of the labels associated with more efficient codes is then proposed to efficiently leverage contextual informations available in the encoder and construct the Most Probable Modes (MPM) list. The proposed coding scheme enables to increase the BD-BR performance in average by 0.09% for the same level of complexity compared to the JEM. Kevin Reuze, Wassim Hamidouche, Pierrick Philippe, Olivier Déforges |
DCC | 4 |
| 2019 | Laboratory and Crowdsourcing Studies of Lip Sync Effect on the Audio-Video Quality Assessment for Videoconferencing ApplicationabstractLip sync is one of the factors that impacts a lot the quality of the videoconferencing experience. In this paper we study the end-user perception of the asynchrony and we try to determine the annoyance threshold of lip-synch error. We are also interested in investigating the mutual interaction between the asynchrony annoyance and changes in the video coding bit rate, spatial video resolution, video IP packet loss, and audio IP packet loss. We conducted two subjective tests in two different environments: laboratory and crowdsourcing. The experimental results show that the audio-video desynchronization annoyance is not an independent factor, but influenced by the video and audio quality. Furthermore, by comparing the results of the two tested methodologies, we show that the crowdsourcing methodology for the quality assessment might be used for audiovisual, video and asynchrony perception assessment, but there is some challenges to consider for the audio quality test. Inès Saidi, Lu Zhang 0037, Vincent Barriac, Olivier Déforges |
ICIP | 4 |
| 2019 | Hardware-friendly DST-VII/DCT-VIII approximations for the Versatile Video Coding StandardabstractVersatile Video Coding (VVC) is the next generation video coding standard expected by the end of 2020. The new concept of Multiple-Transform Selection (MTS) has been introduced in VVC. MTS enables the VVC encoder to select the transform that minimizes the rate-distortion cost among a set of pre-defined trigonometric transforms including the well known Discrete Cosine Transform (DCT)-II, DCT-VIII and Discrete Sine Transform (DST)-VII. Unlike the DCT-II that has fast computing algorithms, the DST-VII and DCT-VIII rely on more complex matrix multiplication.This paper tackles the problem of DST-VII and DCT-VIII approximations based on the DCT-II and an adjustment stage. This latter consists in a multiplication by a band-matrix with low number of non-zero coefficients per row. The approximation problem is first modeled as a constrained integer optimization problem minimizing both error and orthogonality. The genetic algorithm is then used to solve the optimization problem and find the adjustment band-matrix that minimizes a trade-off between error and orthogonality. The proposed solution enables to preserve the coding gain achieved by the MTS and considerably reduces the complexity in terms of required number of multiplications by coefficient. Moreover, the proposed approach is hardwarefriendly and will provide a lightweight shared hardware module for DST-II, DST-VII and DCT-VIII transforms. Wassim Hamidouche, Pierrick Philippe, Camar-Eddine Mohamed, Ahmed Kammoun, Daniel Ménard, Olivier Déforges |
PCS | 6 |
| 2019 | Compression Performance of the Versatile Video Coding: HD and UHD Visual Quality MonitoringabstractVideo compression and content quality have become one of the most research topic in the recent years. Predominantly, trends obviously signpost that the video usage over the Internet is on the upsurge. Simultaneously, users' requirement for enlarged resolution and higher quality is rising. Consequently, a huge effort has been made for video coding technologies and quality monitoring. In this paper, we present a subjective-based comparison as well as an objective measurement between the newest Versatile Video Coding (VVC) and the well-known High Efficiency Video Coding (HEVC) standards. Several videos of various content are selected as tested sequences. Both High Definition (HD) and Ultra High Definition (UHD) resolutions are used in this experiment. An extensive range of bit-rates from low to high bit-rates were selected. These sequences are encoded using both HEVC reference software (HM-16.2) and the latest reference software of VVC (VTM-5.0). Obtained results have shown that VVC outperforms consistently HEVC, for realistic bit rates and quality levels, in the range of 40% on the subjective scale. For the objective measurements, using PSNR, SSIM and VMAF as quality metrics, the quality enhancement of VVC over HEVC is ranging from 31% to 40%, depending on video content and spatial resolution. Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Pierrick Philippe, Jérôme Fournier |
PCS | 3 |
| 2019 | Perceptual Evaluation of Adversarial Attacks for CNN-based Image ClassificationabstractDeep neural networks (DNNs) have recently achieved state-of-the-art performance and provide significant progress in many machine learning tasks, such as image classification, speech processing, natural language processing, etc. However, recent studies have shown that DNNs are vulnerable to adversarial attacks. For instance, in the image classification domain, adding small imperceptible perturbations to the input image is sufficient to fool the DNN and to cause misclassification. The perturbed image, called adversarial example, should be visually as close as possible to the original image. However, all the works proposed in the literature for generating adversarial examples have used the Lpnorms (L0, L2and L∞) as distance metrics to quantify the similarity between the original image and the adversarial example. Nonetheless, the Lpnorms do not correlate with human judgment, making them not suitable to reliably assess the perceptual similarity/fidelity of adversarial examples. In this paper, we present a database for visual fidelity assessment of adversarial examples. We describe the creation of the database and evaluate the performance of fifteen state-of-the-art full-reference (FR) image fidelity assessment metrics that could substitute Lpnorms. The database as well as subjective scores are publicly available to help designing new metrics for adversarial examples and to facilitate future research works. Sid Ahmed Fezza, Yassine Bakhti, Wassim Hamidouche, Olivier Déforges |
QoMEX | 4 |
| 2019 | Visual Security Assessment of Selective Video EncryptionabstractGiven the wide use of videos in various applications and across different devices, this raises the question of their security and confidentiality. In the last decade, many video encryption methods have been proposed in the literature. Accordingly, it becomes necessary to have a reliable assessment tool allowing evaluation of the efficiency of these video encryption methods, especially from the visual security point of view. Usually, the visual security is evaluated through the classical objective signal-based metrics. However, these metrics showed their limits as visual security metric, since they are not designed to deal with the security requirements, such as the determination of content intelligibility. Despite its obvious importance, very few visual security metrics have been proposed for the assessment of video encryption methods. This is mainly due to the lack of ground truth with subjective human scores for video encryption applications. In this paper, we present a new database for visual security assessment of selective video encryption. The database including unencrypted and encrypted video contents generated using different selective encryption schemes, as well as subjective scores, is publicly available to help designing new visual security metrics1. Sid Ahmed Fezza, Wassim Hamidouche, Reda Abdellah Kamraoui, Olivier Déforges |
QoMEX | 4 |
| 2019 | An Adaptive Quantizer for High Dynamic Range Content: Application to Video CodingabstractIn this paper, we propose an adaptive perceptual quantization method to convert the representation of high dynamic range (HDR) content from the floating point data type to integer, which is compatible with the current image/video coding and display systems. The proposed method considers the luminance distribution of the HDR content, as well as the detectable contrast threshold of the human visual system, in order to preserve more contrast information than the perceptual quantizer (PQ) in integer representation. Aiming to demonstrate the effectiveness of this quantizer for HDR video compression, we implemented it in a mapping function on the top of the HDR video coding system based on high efficiency video coding standard. Moreover, a comparison function is also introduced to decrease the additional bit-rate of side information, generated by the mapping function. Objective quality measurements and subjective tests have been conducted in order to evaluate the quality of the reconstructed HDR videos. Subjective test results have shown that the proposed method can improve, in a significant manner, the perceived quality of some reconstructed HDR videos. In the objective assessment, the proposed method achieves improvements over PQ in terms of the average bit-rate gain for metrics used in the measurement. Yi Liu 0004, Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Giuseppe Valenzise, Emin Zerman |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | A Benchmark of DIBR Synthesized View Quality Assessment Metrics on a New Database for Immersive Media ApplicationsabstractDepth-image-based rendering (DIBR) is a fundamental technology in several 3-D-related applications, such as free viewpoint video, virtual reality, and augmented reality. However, new challenges have also been brought in assessing the quality of DIBR-synthesized views since this process induces some new types of distortions, which are inherently different from the distortion caused by video coding. In this paper, we present a new DIBR-synthesized image database with the associated subjective scores. We also test the performances of the state-of-the-art objective quality metrics on this database. This paper focuses on the distortions only induced by different DIBR synthesis methods. Seven state-of-the-art DIBR algorithms, including inter-view synthesis and single-view-based synthesis methods, are considered in this database. The quality of synthesized views was assessed subjectively by 41 observers and objectively using 14 state-of-the-art objective metrics. Subjective test results show that the interview synthesis methods, having more input information, significantly outperform the single-view-based ones. Correlation results between the tested objective metrics and the subjective scores on this database reveal that further studies are still needed for a better objective quality metric dedicated to the DIBR-synthesized views. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
IEEE Trans. Multim. | 4 |
| 2018 | Light Field Image Compression Based on Convolutional Neural Networks and Linear ApproximationabstractComputer vision applications such as refocusing, segmentation and classification become one of the most advanced imaging services. Light Field (LF) imaging systems provide a rich semantic information of the scene. Using a dense set of cameras and microlens arrays (Plenoptic camera), the direction of each ray coming from the scene toward the LF capture system can be extracted and represented by spatial and angular coordinates. However, such imaging system induces many drawbacks including the large amount of data produced and complexity increase for scene representation. In this paper, we propose an efficient LF image coding scheme. This scheme first encodes a sparse set of views using the latest hybrid video encoder (JEM). Then, it estimates a second sparse set of views using a linear approximation. At the decoder side, we use a Deep Learning (DL) approach to estimate the whole LF image from the reconstructed sparse sets of views. Experimental results show that the proposed scheme provides higher visual quality and overcomes the state of the art LF image compression solution by 30 % bitrate gain. Nader Bakir, Wassim Hamidouche, Olivier Déforges, Khouloud Samrouth |
ICIP | 3 |
| 2018 | Live Demonstration: End-to-End Real-Time ROI-based Encryption in HEVC VideosabstractThis paper presents a demonstration setup for live HEVC video coding with Region of Interest (ROI) encryption. The showcased approach splits video frames into independent HEVC tiles and encrypts those belonging to the ROI. This end-to-end content protection scheme is put into practice by integrating the algorithms of selective encryption into Kvazaar HEVC encoder and decryption into openHEVC decoder. The shown implementation performs secure encryption of the ROI in real time with small bit rate and complexity overhead. Naty Ould Sidaty, Marko Viitanen, Wassim Hamidouche, Jarno Vanne, Olivier Déforges |
ISCAS | 5 |
| 2018 | Evaluation of No-reference quality metrics for Ultrasound liver imagesabstractAlthough assessing post-processed medical images is still done by radiologists (rather than computers), numerous algorithms dedicated to medical image processing are developed without taking into consideration the expert's perceived quality scores. In order to evaluate these algorithms, we study in this paper four No-Reference(NR) quality assessment metrics in terms of correlation with perceived scores of experts. These scores were obtained through subjective tests conducted on ultrasound (US) livers images. Results show that one NR metric among the four evaluated performs the best for assessing the quality of US images. However, further study is needed for the development of more suitable NR metrics. Meriem Outtas, Lu Zhang 0037, Olivier Déforges, Wassim Hamidouche, Amina Serir |
QoMEX | 3 |
| 2018 | SC-IQA: Shift compensation based image quality assessment for DIBR-synthesized viewsabstractDepth-image-based-rendering (DIBR) has been used to generate the virtual views for Multi-view videos and Free-viewpoint videos. However, the quality assessment of DIBR-synthesized views is very challenging owing to the new types of distortions induced by inaccurate depth maps, dis-occlusions and image inpainting methods. There exist a large number of object shifts and geometric distortions in the synthesized view which the traditional 2D quality metrics may fail to assess. In this paper, we propose a shift compensation based image quality assessment metric (SC-IQA) for DIBR-synthesized views. Firstly, the global geometric shift is compensated roughly by an SURF + RANSAC homography approach. Then, a multi-resolution block matching method, which performs a more accurate matching, is used to precisely compensate the shift and penalize the local geometric distortion as well. In addition, a visual saliency map is also used as a weighting function. To calculate the final overall quality scores, only the worst blocks are utilized since the biggest distortions have the most effects on the overall perceptual quality. The results show that the proposed metric significantly outperforms the state-of-the-art synthesized view dedicated metrics and the conventional 2D IQA metrics. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
VCIP | 4 |
| 2018 | Cryptanalyzing an image encryption scheme using reverse 2-dimensional chaotic map and dependent diffusion
Mousa Farajallah, Safwan El Assad, Olivier Déforges |
Multim. Tools Appl. | 3 |
| 2018 | NIQSV+: A No-Reference Synthesized View Quality Assessment MetricabstractBenefiting from multi-view video plus depth and depth-image-based-rendering technologies, only limited views of a real 3-D scene need to be captured, compressed, and transmitted. However, the quality assessment of synthesized views is very challenging, since some new types of distortions, which are inherently different from the texture coding errors, are inevitably produced by view synthesis and depth map compression, and the corresponding original views (reference views) are usually not available. Thus the full-reference quality metrics cannot be used for synthesized views. In this paper, we propose a novel no-reference image quality assessment method for 3-D synthesized views (called NIQSV+). This blind metric can evaluate the quality of synthesized views by measuring the typical synthesis distortions: blurry regions, black holes, and stretching, with access to neither the reference image nor the depth map. To evaluate the performance of the proposed method, we compare it with four full-reference 3-D (synthesized view dedicated) metrics, five full-reference 2-D metrics, and three no-reference 2-D metrics. In terms of their correlations with subjective scores, our experimental results show that the proposed no-reference metric approaches the best of the state-of-the-art full reference and no-reference 3-D metrics; and outperforms the widely used no-reference and full-reference 2-D metrics significantly. In terms of its approximation of human ranking, the proposed metric achieves the best performance in the experimental test. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
IEEE Trans. Image Process. | 4 |
| 2017 | Cluster Adapted Signalling for Intra Prediction in HEVCabstractThe High Efficiency Video Coding (HEVC) standard defines 35 Intra Prediction Modes (IPM) to provide an efficient compression of intra coded blocks. Those IPMs are signalled to the decoder through the use of three compression tools: prediction, clustering and coding. In this paper we provide improvements to these three tools through: new labels for the prediction, new tests for the clustering and new coding schemes for the coding. The most significant improvement consists in the provision of a cluster-dependent code: adapting the coding scheme to the available information enables the average symbol cost to get within close margin of the entropy of the data. The system providing the best compression efficiency based on these improvements is then computed, enabling significant reduction in the average cost required to code the IPMs. The proposed method builds a new coding system with the same complexity as HEVC with 0.41% bit-rates savings in All Intra coding configuration. Kevin Reuze, Pierrick Philippe, Wassim Hamidouche, Olivier Déforges |
DCC | 4 |
| 2017 | Real-time and parallel SHVC hybrid codec AVC to HEVC decoderabstractScalable High efficiency Video Coding (SHVC) is the scalable extension of the latest video coding standard High Efficiency Video Coding (HEVC). One of the key novelties introduced by SHVC is that it enables hybrid codec scalability. This basically means that the video layers can be encoded with different video standards providing backward compatibility between codecs. In this paper, we propose a software parallel SHVC decoder in hybrid codec scalability configuration. The proposed design consists of an Advanced Video Coding (AVC) decoder for the Base Layer (BL) and a HEVC decoder for the Enhanced Layer (EL). In order to perform Inter Layer Prediction (ILP), a communication of decoding states and outputs is established between the two decoders. While the native frame based parallelism is still allowed within the two decoders, the proposed design also enables the use of frame based parallelism between the two decoders. The proposed software design enables a real time decoding of the HEVC EL at 2160p60 while the AVC base layer is decoded at 1080p60 for ×2 spatial scalability. Pierre-Loup Cabarat, Wassim Hamidouche, Olivier Déforges |
ICASSP | 3 |
| 2017 | A new perceptual assessment methodology for selective HEVC video encryptionabstractVideo data security is one of the most research topic in the recent years. It is widely used in the multimedia applications such as video-conferencing, Video on Demand and Pay-TV services. Although many video encryption methods and objective measurements have been employed, few real time schemes and no subjective studies have been proposed. In this paper we investigate a set of selective video encryption schemes by encrypting only a few parameters in HEVC video streams. Firstly, we carried out an in-depth subjective study of three proposed selective HEVC video encryption schemes. A panel of observers has participated in this test campaign in order to evaluate the degree of visibility of the encrypted videos at different bitrates. Experimental results are presented and analysed, showing therefore that two proposed selected encryption schemes allow a high perceptual security level by masking the whole details of the video content, while the third scheme achieves a high security level, with a content nearly unidentifiable. In addition, subjective scores can be used as ground truth for assessing selective video encryption methods, instead of classical objective signal-based metrics, which are not correlated with human judgment. Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges |
ICASSP | 3 |
| 2017 | NIQSV: A no reference image quality assessment metric for 3D synthesized viewsabstractThe popularity of 3D applications, such as Free View-point TV (FTV) and Multi-view Video plus Depth (MVD), induces a heavy requirement of synthesized views. However, the quality assessment of synthesized views is very challenging because the corresponding original views (reference views) are usually not available at both encoder and decoder sides. In this paper, we propose a new no-reference quality assessment model to evaluate the quality of 3D synthesized views, called NIQSV (No-reference Image Quality assessment of Synthesized Views). This metric is based on the hypothesis that a good quality image is composed of flat areas (objects) separated by sharp edges, and the quality estimation involves only a set of simple morphological operators. NIQSV integrates the distortions of all the components, and then uses an edge image to weight the final distortions since the distortions of synthesized views mainly happen around object edges. The experimental results show that the proposed metric outperforms traditional 2D metrics and ranks among the best of dedicated 3D synthesized and full reference metrics. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
ICASSP | 4 |
| 2017 | An adaptive perceptual quantization method for HDR video codingabstractThis paper presents a new adaptive perceptual quantization method for the High Dynamic Range (HDR) content. This method considers the luminance distribution of the HDR image as well as the Minimum Detectable Contrast (MDC) thresholds to preserve the contrast information during quantization. Base on this method, we develop a mapping function for HDR video compression and apply it to a HEVC Main 10 Profile-based video coding chain. Our experiments show that the proposed mapping function can efficiently improve the quality of the reconstructed HDR video in both objective and subjective assessments. Yi Liu 0004, Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Giuseppe Valenzise, Emin Zerman |
ICIP | 4 |
| 2017 | Multi-output speckle reduction filter for ultrasound medical images based on multiplicative multiresolution decompositionabstractUltrasonographic examination, either as visual inspection or quantitative analysis, is less effective than other medical imaging systems due to speckle noise. The state-of-the-art speckle reduction methods often offers an effective speckle reduction but generally they suffer from oversmoothig, blurring effect and man-made/artificial appearance. In this paper, a new Multi-Output Filter based on a Multiplicative Multiresolution Decomposition (MOF-MMD) is proposed. This multiscale based method, particularly efficient in the case of multiplicative noise, enhances distinctively three outputs: edges, texture and the global image. The multi-output filter aims at offering an enhanced images according to the features desired by radiologists. The different structures, textures and edges are filtered according to the contour image obtained by morphological operators. Finally, we compare the MOF-MMD method with two state-of-the-art speckle reduction methods in terms of speckle reduction capacity and image quality improvement. The results show that the proposed method offers an effective speckle reduction with an improvement of the image quality without blurry and over-smoothing effect. Meriem Outtas, Lu Zhang 0037, Olivier Déforges, Amina Serir, Wassim Hamidouche |
ICIP | 3 |
| 2017 | Compression efficiency of the emerging video coding toolsabstractWith the drastic increasing of multimedia applications and video coarse consumption, video compression and content quality evaluation have become an exciting and challenging topic. Recently, a new coding tool has been developed under the Joint Exploration Model (JEM) software with the main goal to provide a high bit rate saving compared to the HEVC standard. In this paper we present a performance-based comparison between the JEM and HEVC reference software (HM) through an objective measurements and a subjective quality assessments. A set of video sequences, in two spatial resolutions High Definition (HD) and Ultra-High Definition (UHD), have been used in this study. These videos are encoded using both JEM and HM software at different bitrates. Results have shown that the JEM codec enables, subjectively, a quality enhancement up to 40% at similar low bit rates. Objectively, this quality improvement is ranging from 35% to 37% depending on the spatial resolution. However, at high bit rate, the HM reference software enables a high video quality and thus its becomes more difficult to perceive the quality enhancement is about brought by the JEM codec. In addition, some video contents are difficult to encode and, consequently, the JEM enables only slight perceived quality improvement especially at the highest considered bitrates and for 4K resolutions. Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Pierrick Philippe |
ICIP | 3 |
| 2017 | Evaluation of single-artifact based video quality metrics in video communication contextabstractFor an accurate assessment of media quality, it is essential not only to compute an overall quality measure, but also to identify the type of occurring distortions. In this paper, we focus on a set of no-reference single-artifact based metrics developed by the MOAVI project. We carried out a correlation analysis in order to evaluate the performance of these metrics on three databases with a large sample of distortion types. This study will be used for setting up a video quality monitoring tool box. Inès Saidi, Lu Zhang 0037, Vincent Barriac, Olivier Déforges |
QoMEX | 4 |
| 2017 | Emerging video coding performance: 4K quality monitoringabstractThe new coding tools, developed under the Joint Exploration Model (JEM) software, have been proposed with the main goal to explore their potential coding gain in the perspective to develop a new video coding standard. In this paper we present a performance-based comparison between the JEM and HEVC reference software (HM) through a set of subjective quality assessments. Different video sequences, encoded using both JEM and HM software at different bitrates, have been used in this experiment. Results have shown that the JEM codec enables, subjectively, a quality enhancement up to 40% at similar low bit rates. Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Pierrick Philippe |
QoMEX | 3 |
| 2017 | Real-time selective video encryption based on the chaos system in scalable HEVC extension
Wassim Hamidouche, Mousa Farajallah, Naty Ould Sidaty, Safwan El Assad, Olivier Déforges |
Signal Process. Image Commun. | 5 |
| 2016 | Optimal Bitrate Allocation for High Dynamic Range and Wide Color Gamut Services Deployment Using SHVCabstractThe scalable video coding enables to compress video contents into a hierarchical layered representation, each layer depicts an enhanced version of the underlying layer. SHVC is the scalable extension of HEVC and enables spatial, SNR, color-gamut, codec and bitdepth scalability. It has been proved, in the MPEG investigations prior to the recent Call for Evidence, that SHVC can support SDR-to-HDR scalability by using the color gamut scalability, when SDR and HDR signals are placed in different color gamuts. This way, SHVC can be used to address future backward compatible issues in the HDR and WCG services deployment. In this paper, we exploit the impact of bitrate ratio over performance in scalable schemes to design an adaptive rate control algorithm suitable for such deployment, considering adjustable quality and bandwidth constraints. Our method dynamically adjusts the bitrate ratio between two layers during encoding in the most quality-related optimal way under specified constraints. The proposed method is tested on scalable combinations of HD/UHD, R.709/DCI-P3/R.2020 and SDR/HDR video contents, and reduces the average overhead introduced by SHVC compared to the single-layer HEVC encoding by 23%. Thibaud Biatek, Wassim Hamidouche, Jean-François Travers, Olivier Déforges |
DCC | 4 |
| 2016 | Low complexity transform competition for HEVCabstractThe use of multiple transforms in video coding can lead to substantial bit-rate savings. However, these savings come at the expense of increased coding complexity and storage requirements, which challenge the usability of this approach. In this paper, a systematic procedure is proposed to design low complexity systems making use of transform competition. Multiple trade-offs accommodating the complexity are unveiled and it is demonstrated that they can keep a certain level of performance. Compared to the HEVC standard, some of them provide bit-rate savings around 2% with a 50% increase in the encoding time, using less than 4 kB of extra ROM and no added decoding complexity. Adria Arrufat, Pierrick Philippe, Kevin Reuze, Olivier Déforges |
ICASSP | 4 |
| 2016 | Adaptive rate control algorithm for SHVC: Application to HD/UHDabstractScalable video coding consists in compressing the video sequence into a layered bitstream where each layer refers to different spatial, temporal or quality representation of the video. Scalability enables compression gain compared to the simulcast encoding of layers thanks to inter-layer predictions. The scalable HEVC extension (SHVC) is the latest scalable technology promising up to 30% bitrate gains under the common test conditions, defined by JCT-VC. These conditions do not consider UHD and use fixed quantization step, which is not relevant in operational environment. In this paper, we propose an innovative adaptive rate control algorithm for SHVC. We consider HD as a base layer and UHD as an enhancement layer, with a constant global bitrate and a dynamic bitrate ratio adjustment between layers. The proposed algorithm is evaluated on a UHD data set where enables on average a BD-BR gain of 4.25% compared to a fixed-ratio encoding. Thibaud Biatek, Wassim Hamidouche, Jean-François Travers, Olivier Déforges |
ICASSP | 4 |
| 2016 | Audiovisual quality study for videoconferencing on IP networksabstractIn this paper, an audiovisual quality assessment experiment was conducted on audiovisual clips collected using a PC-based videoconferencing application connected via a local IP network. The analyses of experimental results provided a better understanding of the influence of network impairments (packet loss, jitter, delay) on perceived audio and video qualities, as well as their interaction effect on the overall audiovisual quality in videoconferencing applications. We updated the human perception acceptability limits of audio-video synchronization for video conferencing. Further, we investigated the contribution of this synchronization to the audiovisual quality independently and accompanied with network impairments. Finally, we proposed an integration model to estimate the audiovisual quality in the studied context. Inès Saidi, Lu Zhang 0037, Vincent Barriac, Olivier Déforges |
MMSP | 4 |
| 2016 | Pre-encoding based statistical-multiplexing for hybrid delivery of UHD services using SHVCabstractThe scalable video coding consists in encoding the video content into multiple representations, called layers, where each one refers to a specific version of the content. The scalable extension of the High Efficiency Video Coding standard SHVC is currently considered in ATSC and DVB to carry out layered and scalable programs which enables to target multiple equipments, and is also used to ensure backward compatibility with legacy receivers. In the case of hybrid delivery of HD and UHD services using SHVC, these services are encoded in two layers including base and enhancement layers, which are then broadcasted over separated channels. In this paper, a statistical multiplexing method is proposed for broadcasting of UHD services in this hybrid scenario. This innovative method considers both variable bitrate among programs and optimal SHVC layers coding, which was not considered in the existing approaches. The proposed method enables to reduce the overhead introduced by SHVC compared to the single-layer encoding by 3.3% in average while maintaining smooth quality variations among programs. Thibaud Biatek, Wassim Hamidouche, Jean-François Travers, Olivier Déforges |
PCS | 4 |
| 2016 | Intra prediction modes signalling in HEVCabstractThe High Efficiency Video Coding (HEVC) standard defines 35 Intra Prediction Modes (IPM) to provide an efficient compression of intra coded blocks. To signal these IPMs to the decoder a list of 3 Most Probable Modes (MPM) is created based on the IPMs of the neighbour Intra coded blocks. These MPMs will be transmitted in the bit stream with a reduced number of bits. However, the signalling scheme used in HEVC is not optimal and can be further improved. In this paper a method is proposed to enhance the Intra mode signalling in HEVC. The IPM signaling is first modelled as a tree with forks and leaves representing tests and labels, respectively. The proposed solution introduces new decision tree process by adding new tests and new labels not considered in HEVC. This solution provides a systematic way to find the best signalling scheme for a given set of data. Experimental results show that the proposed solution enables to reduce the BD-Rate by 0.38% in all Intra coding configuration. Kevin Reuze, Pierrick Philippe, Olivier Déforges, Wassim Hamidouche |
PCS | 3 |
| 2016 | Interactive vs. non-interactive subjective evaluation of IP network impairments on audiovisual quality in videoconferencing contextabstractIn this paper, we present a subjective audiovisual quality assessment experiment realized using a PC based video-conferencing application connected via a local IP network. The experiment was conducted under two different scenarios: a non-interactive and an interactive conversational one. We present the effects of network impairments (packet loss, delay) on perceived audiovisual, audio and video quality. We evaluate the impact of scene complexity on the quality perception in case of video calls. We establish a comparison between the perception of multimedia quality in interactive and non-interactive context. The results presented in this paper show a dependency between the perceived quality and the scene complexity: the perceived quality of a high spatial and temporal complex scene is inferior to that of the less complex one. In addition, our findings show that the audio-video synchronization acceptability do not differ between the interactive and the non-interactive experimental context. Inès Saidi, Lu Zhang 0037, Vincent Barriac, Olivier Déforges |
QoMEX | 4 |
| 2016 | 4K Real-Time and Parallel Software Video Decoder for Multilayer HEVC ExtensionsabstractTwo High Efficiency Video Coding (HEVC) extensions, namely, the scalable HEVC (SHVC) extension and multiview HEVC (MV-HEVC) extension, have been finalized in July 2014 by the Moving Picture Experts Group and Video Coding Experts Group. These two extensions enable additional features not covered in the first version of the HEVC standard such as spatial, fidelity, bitdepth, and color gamut scalability, as well as stereoscopic and multiview representations. In this paper, we propose a software parallel decoder architecture for the HEVC standard and its multilayer extensions, including SHVC and MV-HEVC extensions. The decoder consists of multiple instances of the OpenHEVC decoder, one instance to decode each layer with a communication between dependent layers to perform inter-layer predictions. The proposed multilayer HEVC decoder is parallel friendly and supports both wavefront parallelism to simultaneously process adjacent rows of the frame and frame-based parallelism to decode a set of temporal and spatial frames in parallel. Moreover, the most time-consuming operation introduced in the SHVC extension, namely, the resampling of the inter-layer reference picture in spatial scalability, is optimized in single instruction multiple data for x86 platform. We assess the complexity of the multilayer HEVC decoder with respect to the simulcast configuration. The multilayer decoder decoding two SHVC layers introduces in average 40%-71% additional complexity compared with the single layer HEVC decoder. Moreover, the low level optimizations with a hybrid parallel processing solution enable a real-time decoding of 4Kp60 enhancement layer on a 6-core Intel i7 processor running at 3.4 GHz. Wassim Hamidouche, Mickaël Raulet, Olivier Déforges |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | LAR-LLC: A Low-Complexity Multiresolution Lossless Image CodecabstractThis paper presents a new scalable locally adaptive resolution lossless low-complexity (LAR-LLC) image codec. It is based on the LAR framework that is a multiresolution compression method supporting both lossy and lossless coding. To achieve an efficient low-complexity solution, each processing stage of the LAR is modified. For the first step, consisting of a pyramidal decomposition, a new reversible transform called hierarchical diagonal$S$transform (HD-ST) is proposed. The HD-ST operates on sets of data pairs, requiring only shift and add/sub operations. The second step performs the prediction of the transformed coefficients. The prediction scheme considers both inter- and intra-level information, and involves fixed weights. Then, a classification process is introduced to separate prediction errors into subclasses, using a context modeling approach. Finally, each subclass is coded by the Huffman coding algorithm. The results of the lossless compression experiments showed that LAR-LLC achieves the same compression performance as JPEG2000 with a lower complexity. Yi Liu 0004, Olivier Déforges, Khouloud Samrouth |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Selective video encryption using chaotic system in the SHVC extensionabstractIn this paper we investigate a selective video encryption in the scalable HEVC extension (SHVC). The SHVC extension encodes the video in several layers corresponding to different spatial and quality representations of the video. We propose a selective encryption solution using a chaotic-based encryption system. The proposed solution encrypts a set of sensitive parameters with a minimum complexity overhead, at constant bitrate and SHVC format compliant. Experimental results compare the performance of three encryption schemes: encrypt only the lowest layer, all layers, and only the highest layer. The first two schemes achieve a high security level with a drastic degradation in the decoded video, while the last scheme enables a perceptual video encryption by decreasing the quality of the highest layer below the quality of the clear layers. Wassim Hamidouche, Mousa Farajallah, Mickaël Raulet, Olivier Déforges, Safwan El Assad |
ICASSP | 4 |
| 2015 | Mode-dependent transform competition for HEVCabstractTransform coding plays a key role in state-of-the-art video coders, such as HEVC. However, transforms used in current solutions do not cover the varieties of video coding signals. This work presents an adaptive transform design method that enables the use of multiple transforms in HEVC. A different transform set is learnt for each intra prediction mode, allowing the video encoder to perform better decisions regarding block sizes, prediction modes and transforms. Different systems are proposed to accommodate trade-offs between complexity and performance. Bit rate reductions in the range of 2% to 7% are reported, depending on complexity. Adria Arrufat, Pierrick Philippe, Olivier Déforges |
ICIP | 3 |
| 2015 | ROI encryption for the HEVC coded video contentsabstractIn this paper we investigate privacy protection for the HEVC standard based on the tile concept. Tiles in HEVC enable the video to be split into independent rectangular regions. Two solutions are proposed to encrypt the tiles containing the Region Of Interest (ROI). The first solution performs encryption at the bitstream level by encrypting all HEVC syntax elements within the ROI tiles. The second solution enables a selective encryption of the ROI tiles under constant bitrate and format compliant requirements. To avoid temporal propagation of the encryption outside the ROI boundaries caused by inter prediction, the motion vectors of non ROI regions are restricted inside the non encrypted tiles in the reference frames. Simulation results show that the proposed solutions perform secure and adaptive encryption of ROI in the HEVC video. Moreover, the bitrate overhead caused by the MVs restriction window varies between 1%-2.5% depending on both the video content and the number of tiles within the frame. Mousa Farajallah, Wassim Hamidouche, Olivier Déforges, Safwan El Assad |
ICIP | 3 |
| 2014 | Multi-core software architecture for the scalable HEVC decoderabstractThe scalable high efficiency video coding (SHVC) standard aims to provide features of temporal, spatial and quality scalability. In this paper we investigate a pipeline and parallel software architecture for the SHVC decoder. The proposed architecture is based on the OpenHEVC software which implements the high efficiency video coding (HEVC) decoder. The architecture of the SHVC decoder enables two levels of parallelism. The first level decodes the base layer and the enhancement layers in parallel. The second level of parallelism performs the decoding of both the base layer and enhancement layers in parallel through the HEVC high level parallel processing solutions, including tile and wavefront. Up to the best of our knowledge, it is the first real time and parallel software implementation of the SHVC decoder. On an Intel Xeon processor running at 3.2 GHz, the SHVC decoder reaches the decoding of 1600p enhancement layer at 40 fps for x1.5 spatial scalability with using six concurent threads. Wassim Hamidouche, Mickaël Raulet, Olivier Déforges |
ICASSP | 3 |
| 2014 | Real time SHVC decoder: Implementation and complexity analysisabstractThe Scalable High efficiency Video Coding (SHVC) standard is developed to offer spatial and quality scalability with high coding efficiency. In this paper we investigate a complexity analysis of a real time and parallel SHVC decoder. We first provide details on the implementation of the SHVC decoder including its low level optimizations. Furthermore, we introduce parallelism tools integrated in the SHVC software for parallel decoding. These tools include frame-based parallelism to decode a set of temporal and spatial frames in parallel as well as wavefront parallelism to simultaneously process separated regions of a picture. We assessed through experimental results the complexity of the real time SHVC decoder in different coding configurations. The SHVC decoder with two layers introduces in average an additional complexity of 43 to 80% in respect to a simulcast configuration. The low level optimizations together with a hybrid parallelism solution enables a real time decoding of 1600p40 enhancement layer on an Intel i7 processor. Wassim Hamidouche, Mickaël Raulet, Olivier Déforges |
ICIP | 3 |
| 2014 | Parallel SHVC decoder: Implementation and analysisabstractThe new Scalable High efficiency Video Coding (SHVC) standard is based on a multi-loop coding structure which requires the total decoding of all intermediate layers. The decoding complexity becomes then a real issue, especially for a real time decoding of ultra high video resolutions. A parallel processing architecture is proposed to reduce both the decoding time and the latency of the SHVC decoder. The proposed solution combines the high level parallel processing solutions defined in the HEVC standard with an extension of the frame-based parallelism. The latter solution enables the decoding of several spatial and temporal SHVC frames in parallel to enhance both decoding frame rate and latency. The wavefront parallel processing solution is used for more coarse level of granularity. The proposed hybrid parallel processing approach achieves a near optimal speedup and provides a good trade-off between decoding time, latency and memory usage. On a 6 cores Xeon processor, the parallel SHVC decoder performs a real time decoding of 1600p60 video resolution. Wassim Hamidouche, Mickaël Raulet, Olivier Déforges |
ICME | 3 |
| 2014 | Non-separable mode dependent transforms for intra coding in HEVCabstractTransform coding plays a crucial role in video coders. Recently, additional transforms based on the DST and the DCT have been included in the latest video coding standard, HEVC. Those transforms were introduced after a thoroughly analysis of the video signal properties. In this paper, we design additional transforms by using an alternative learning approach. The appropriateness of the design over the classical KLT learning is also shown. Subsequently, the additional designed transforms are applied to the latest HEVC scheme. Results show that coding performance is improved compared to the standard. Additional results show that the coding performance can be significantly further improved by using non-separable transforms. Bitrate reductions in the range of 2% over HEVC are achieved with those proposed transforms. Adria Arrufat, Pierrick Philippe, Olivier Déforges |
VCIP | 3 |
| 2014 | Rate-distortion optimised transform competition for intra coding in HEVCabstractState of the art video coders are based on prediction and transform coding. The transform decorrelates the signal to achieve high compression levels. In this paper we propose improving the performances of the latest video coding standard, HEVC, by adding a set of rate-distortion optimised transforms (RDOTs). The transform design is based upon a cost function that incorporates a bit rate constraint. These new RDOTs compete against classical HEVC transforms in the rate-distortion optimisation (RDO) loop in the same way as prediction modes and block sizes, providing additional coding possibilities. Reductions in BD-rate of around 2% are demonstrated when making these transforms available in HEVC. Adria Arrufat, Pierrick Philippe, Olivier Déforges |
VCIP | 3 |
| 2014 | A joint 3D image semantic segmentation and scalable coding scheme with ROI approachabstractAlong with the digital evolution, image post-production and indexing have become one of the most advanced and desired services in the lossless 3D image domain. The 3D context provides a significant gain in terms of semantics for scene representation. However, it also induces many drawbacks including monitoring visual degradation of compressed 3D image (especially upon edges), and increased complexity for scene representation. In this paper, we propose a semantic region representation and a scalable coding scheme. First, the semantic region representation scheme is based on a low resolution version of the 3D image. It provides the possibility to segment the image according to a desirable balance between 2D and depth. Second, the scalable coding scheme consists in selecting a number of regions as a Region of Interest (RoI), based on the region representation, in order to be refined at a higher bitrate. Experiments show that the proposed scheme provides a high coherence between texture, depth and regions and ensures an efficient solution to the problems of compression and scene representation in the 3D image domain. Khouloud Samrouth, Olivier Déforges, Yi Liu 0004, Wassim El Falou |
VCIP | 2 |
| 2013 | Efficient Parallelization of Different HEVC Decoding StagesabstractSummary form only given. In this paper we present efficient parallelization implementations for different stages of the HEVC decoder, which are LCU decoding, deblocking filtering and SAO filtering. Each of the stages are parallelized in separate passes. The LCU decoding is parallelized using Wave front Parallel Processing (WPP). Deblocking and SAO filtering are parallelized by segmenting each picture into separate regions of consecutive LCU rows and processing each of the regions in a concurrent fashion. On a 6 core machine with 6 threads running concurrently, experimental results showed an average accelerating factor of 4.6, 5, 5.35 for the LCU decoding stage and 4.5, 4.9, 5 for deblocking filtering stage and 4, 4.5 and 5 for SAO filtering stages on HD, 1600p and 2160p sequences respectively. Anand Meher Kotra, Mickaël Raulet, Olivier Déforges |
DCC | 3 |
| 2013 | Comparison of different parallel implementations for deblocking filter of HEVCabstractUnlike deblocking filter of H.264/AVC, deblocking filter of HEVC is computationally less complex and offers more parallelization possibilities. In this paper we present comparison of three different parallelization implementations of deblocking filter which operate on picture by picture basis. The first method filters the vertical edges and horizontal edges in separate passes. The other two methods combine the vertical and horizontal edge filtering in a single pass. On a 6 core machine with 6 threads running concurrently, experimental results showed an average accelerating factor of 4.5, 5 and 5 for each of the implementation methods on 1080p, 1600p and 2160p sequences respectively. Anand Meher Kotra, Mickaël Raulet, Olivier Déforges |
ICASSP | 3 |
| 2013 | Spatiotemporal texture synthesis and region-based motion compensation for video compression
Fabien Racapé, Olivier Déforges, Marie Babel, Dominique Thoreau |
Signal Process. Image Commun. | 2 |
| 2012 | Automatic generation of synthesizable hardware implementation from high level RVC-cal descriptionabstractData process algorithms are increasing in complexity especially for image and video coding. Therefore, hardware development using directly hardware description languages (HDL) such as VHDL or Verilog is a difficult task. Current research axes in this context are introducing new methodologies to automate the generation of such descriptions. In our work we adopted a high level and target-independent language called CAL (Caltrop Actor Language). This language is associated with a set of tools to easily design dataflow applications and also a hardware compiler to automatically generate the implementation. Before the modifications presented in this paper, the existing CAL hardware back-end did not support some high-level features of the CAL language. Consequently, high-level designed actors had to be manually transformed to be synthesizable. In this paper, we introduce a general automatic transformation of CAL descriptions to make these structures compliant and synthesizable. This transformation analyses the CAL code, detects the target features and makes the required changes to obtain synthesizable code while keeping the same application behavior. This work resolves the main bottleneck of the hardware generation flow from CAL designs. Khaled Jerbi, Mickaël Raulet, Olivier Déforges, Mohamed Abid |
ICASSP | 3 |
| 2011 | Adaptive pixel/patch-based synthesis for texture compressionabstractThis paper presents an adaptive scheme for synthesizing missing textured regions. In synthesis-based compression approaches, large textures are removed at encoder side and filled in at decoder side. This work proposes a synthesizer in which both complementary pixel-based and patch-based approaches are used. According to results shown by synthesis algorithms, patch-based and pixel-based approaches are efficient with different kinds of texture. Two algorithms are adapted to the region synthesis context: the sample patch is built from surrounding texture and potential anchor blocks inside the removed region; the scan order depends on a confidence map in order to take advantage of the designed patch; DCT-descriptors provide the minimum size of matching window which is a required parameter for both synthesizers. Finally an "a posteriori" gradient based assessor enables the synthesizer to switch between algorithms. Fabien Racapé, Simon Lefort, Edouard François, Marie Babel, Olivier Déforges |
ICIP | 5 |
| 2009 | Subjective and objective quality evaluation of lar coded art imagesabstractQuality assessment is of major importance when designing and testing an image/video coding technique. Compression performances are usually evaluated by means of rate-distortion curves. However, the PSNR is commonly employed as the distortion measure. We hereby present a full quality assessment benchmark for the LAR (locally adaptive resolution) coder. We conducted a subjective experiment, where nineteen observers were asked to assess the perceptual quality of LAR coded images under normalized viewing conditions. Furthermore, five objective quality assessment metrics were used in order to determine the most suitable metric for the LAR coder. Finally, both JPEG and JPEG200 images were generated and assessed during the subjective experiment in order to define the optimal quality metric which should be used when comparing the codecs' output images quality. Clément Strauss, François Pasteau, Florent Autrusseau, Marie Babel, Laurent Bédat, Olivier Déforges |
ICME | 6 |
| 2009 | Motion tubes for the representation of image sequencesabstractIn this paper, we introduce a novel way to represent an image sequence, which naturally exhibits the temporal persistence of the textures. Standardized representations have been thoroughly optimized, and getting significant improvements has become more and more difficult. As an alternative, Analysis-Synthesis (AS) coders have focused on the use of texture within a video coder. We introduce here a new AS representation of image sequences that remains close to the classic block-based representation. By tracking textures throughout the sequence, we propose to reconstruct it from a set of moving textures which we call motion tubes. A new motion model is then proposed, which allows for motion field continuities and discontinuities, by hybridizing block matching and a low-computational mesh-based representation. Finally, we propose a bi-predictional framework for motion tubes management. Matthieu Urvoy, Nathalie Cammas, Stéphane Pateux, Olivier Déforges, Marie Babel, Muriel Pressigout |
ICME | 4 |
| 2008 | Software synthesis of CAL actors for the MPEG reconfigurable Video Coding frameworkabstractThe MPEG reconfigurable video coding (RVC) framework aims to provide a unified specification of all video technology. In this framework, a decoder is modularly built as a configuration of video coding tools taken from the MPEG toolbox library. The elements of the library are specified using the CAL actor language. CAL is a dataflow based language providing computation models that are concurrent and modular. This paper presents a synthesis tool that from a CAL specification generates C code. Indeed, code generators are fundamental supports for the deployment and success of the MPEG RVC framework. This paper focuses on the automatic translation of a CAL actor. This approach has been used to obtain a C implementation of the inverse DCT module which is part of the MPEG-4 Simple Profile decoder, chosen by MPEG experts to validate the RVC approach. The generated code is validated against the original CAL description and simulated using the Open Dataflow environment. Ghislain Roquier, Matthieu Wipliez, Mickaël Raulet, Jean-François Nezan, Olivier Déforges |
ICIP | 5 |
| 2008 | Scalable lossless and lossy image coding based on the RWHaT+P pyramid and the inter-coefficient classification methodabstractNext generations of still image codecs should not only have to be efficient in terms of compression ratio, but also propose other functionalities such as scalability, lossy and lossless abilities, region of interest coding, etc. In previous works, we have proposed the LAR compression method covering these requirements. In particular, the RWHaT + P pyramid has recently been presented as a powerful reversible scalable coding technique. This paper introduces new significant improvements by the use of an inter-coefficient classification method. Results are discussed and compared to the state of the art. Olivier Déforges, Marie Babel, Laurent Bédat, Véronique Coat |
ICME | 1 |
| 2008 | Code generation for the MPEG Reconfigurable Video Coding framework: From CAL actions to C functionsabstractThe MPEG reconfigurable video coding (RVC) framework is a new standard under development by MPEG that aims at providing a unified specification of current MPEG video coding technologies. In this framework, a decoder is built as a configuration of video coding modules taken from the standard ldquoMPEG toolbox libraryrdquo. The elements of the library are specified using the CAL actor language (CAL). CAL is a dataflow based language providing computation models that are concurrent and modular. This paper describes a synthesis tool that from a CAL specification automatically generates compilable C-code. Code generators are fundamental supports for the deployment and success of the MPEG RVC framework. This paper focuses on the automatic translation of CAL actions, which is the first step to a complete actor translation. The techniques described here enable to automatically generate C-code according to a finite set of rules. This approach has been used to obtain a C implementation of the IDCT module which is one element of the RVC library. The generated code is validated against the original CAL dataflow program simulated using the open dataflow environment. Matthieu Wipliez, Ghislain Roquier, Mickaël Raulet, Jean-François Nezan, Olivier Déforges |
ICME | 5 |
| 2008 | Rho-Domain for Low-Complexity Rate Control on MPEG-4 Scalable Video CodingabstractScalable Video Coding was designed in response to the growing need for flexibility in video transmission over networks and channels. MPEG-4 Scalable Video Coding (SVC) is a recently finalized standard which introduces new coding tools such as spatial, temporal and quality scalability, to produce a layer-based scalable video stream. Additionally, inter-layer prediction allows a layer to use information from other layers as a basis for motion and texture prediction, improving the overall coding efficiency. Rate control is a capital issue in video coding, as it is designed to regulate the bitrate at the output of the encoder and keep it close to a specified constraint. Whereas rate control has been extensively studied for non-scalable video coding, only few propositions were made for scalable video coding. In this paper, we adapt an attractive rate control approach, based on a bitrate modeling framework called $\rho$-domain, for scalable video coding. We show that this model performs well on all spatial, temporal and quality scalabilities, and handles inter-layer prediction quite accurately. After validating the approach in MPEG-4 SVC, we use the rho-domain model to build a simple-accurate rate control scheme. Results show that the mean frame bitrate error is below 7% on a representative set of configurations, while the impact on the complexity of the encoder is very low. Yohann Pitrey, Yann Serrand, Marie Babel, Olivier Déforges |
ISM | 4 |
| 2008 | A Flexible Heterogeneous Hardware/Software Solution for Real-Time HD H.264 Motion EstimationabstractQuarter-pixel accuracy and variable block-size significantly enhance compression performances of the MPEG-4 AVC/H.264 video compression standard over its predecessors, but also significantly increase computation requirements. Firstly, a digital signal processor (DSP)-based solution that achieves real-time integer motion estimation is proposed. Fractional-pixel refinement is too computationally intensive to be efficiently processed on a software-based processor. To address this restriction, a flexible and low complexity VLSI subpixel refinement coprocessor is designed. Thanks to an improved datapath, a high throughput is achieved with low logic resources. Finally, an heterogeneous (DSP-field-programmable gate array) solution to handle real-time motion estimation with variable block-size and fractional-pixel accuracy for high-definition video is studied. This solution, combining programmability and efficiency, achieves motion estimation of 720 p sequences at up to 60 fps. Fabrice Urban, Ronan Poullaouec, Jean-François Nezan, Olivier Déforges |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2007 | Context-Based Scalable Coding and Representation of High Resolution Art Pictures for Remote Data AccessabstractEROS is the largest database in the world of high resolution art pictures. The TSAR project is designed to open it in a secure, efficient and user-friendly way that involves cryptography and watermarking as well as compression and region-level representation abilities. This paper more particularly addresses the two last points. The LAR codec is first presented as a suitable solution for picture encoding with compression ranging from highly lossy to lossless. Then, we detail the concept of self-extracting region representation, which consists of performing a segmentation process at both the coder and decoder from a highly compressed image, and later locally enhancing the image in a region of interest. The overall scheme provides an efficient, consistent solution for advanced data browsing. Marie Babel, Olivier Déforges, Laurent Bédat, Jean Motsch |
ICME | 2 |
| 2007 | Color LAR Codec: A Color Image Representation and Compression Scheme Based on Local Resolution Adjustment and Self-Extracting Region RepresentationabstractWe present an efficient content-based image coding called locally adaptive resolution (LAR) offering advanced scalability at different semantic levels, i.e., pixel, block, and region. A local analysis of image activity leads to a nonuniform block representation supporting two layers of image description. The first layer provides global information encoded in the spatial domain enabling a low bit rate while preserving contours. The second layer holds texture information encoded in the spectral domain, enabling scalable bitstream in accordance with the required quality. This basic LAR coding leads to an efficient progressive compression, evaluated through subjective quality tests. Its nonuniform block representation also allows a hierarchical region representation providing higher semantic functionalities. More precisely, the segmentation process can be simultaneously performed at both the coder and the decoder from only the luminance component highly compressed by the first coding layer. This solution provides a representation at a region level while avoiding any contour encoding overhead. Region enhancement can then be realized through the second layer. Furthermore, very high compression of the chromatic components is achieved thanks to this region representation. In this scheme, a low-cost chromatic control, which was first introduced during the segmentation process, increases the consistency of region representation in terms of color. Olivier Déforges, Marie Babel, Laurent Bédat, Joseph Ronsin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | Interleaved S+P pyramidal decomposition with refined prediction modelabstractScalability and others functionalities such as the region of interest encoding become essential properties of an efficient image coding scheme. Within the framework of lossless compression techniques, S+P and CALIC represent the state-of-the-art. The proposed interleaved S+P algorithm outperforms these method while providing the desired properties. Based on the LAR (locally adaptive resolution) method, an original pyramidal decomposition combined with a DPCM scheme is elaborated. This solution uses the S-transform in such a manner that a refined prediction context is available for each estimation steps. The image coding is done in two main steps, so that the first one supplies a LAR low-resolution image of good visual quality, and the second one allows a lossless reconstruction. The method exploits an implicit context modelling, intrinsic property of our content-based quad-tree like representation. Marie Babel, Olivier Déforges, Joseph Ronsin |
ICIP (2) | 2 |
| 2003 | Lossless and lossy minimal redundancy pyramidal decomposition for scalable image compression techniqueabstractWe present a new scalable compression technique dealing simultaneously with both lossy and lossless image coding. An original DPCM scheme with refined context is introduced through a pyramidal decomposition adapted to the LAR (locally adaptive resolution) method, which becomes by this way fully progressive. An implicit context modeling of the prediction errors, due to the low resolution image representation including variable block size structure, is then exploited to the for lossless compression purpose. Marie Babel, Olivier Déforges |
ICASSP (3) | 2 |
| 2003 | Rapid prototyping for an optimized MPEG-4 decoder implementation over a parallel heterogenous architectureabstractSequential MPEG-4 solutions actually developed for single processors try to integrate the most functionalities as possible in an unique software, and are generally oversized compared with the actual service requirement. Moreover, they can hardly be projected onto multiprocessors targets, leading to an extra load of source code and calculations, but also to a sub-optimal use of the architecture parallelism. This paper introduces a distributed MPEG-4 application, where the system part is hosted by a standard PC, and the video decoder is supported by a multi-DSPs board. In particular, we present our AVSynDEx methodology allowing both an incremental building, an easy update on the video decoder description, and a quasi-automatic implementation onto a multi-C6x platform. We also define a global scheduler managing the parallel execution of the video and system applications. Nicolas Ventroux, Jean-François Nezan, Mickaël Raulet, Olivier Déforges |
ICASSP (2) | 4 |
| 2003 | Lossless and lossy minimal redundancy pyramidal decomposition for scalable image compression techniqueabstractWe present a new scalable compression technique dealing simultaneously with both lossy and lossless image coding. An original DPCM scheme with refined context is introduced through a pyramidal decomposition adapted to the LAR (locally adaptive resolution) method, which becomes by this way fully progressive. An implicit context modeling of the prediction errors, due to the low resolution image representation including variable block size structure, is then exploited to the for lossless compression purpose. Marie Babel, Olivier Déforges |
ICME | 2 |
| 2003 | Rapid prototyping for an optimized MPEG4 decoder implementation over a parallel heterogeneous architectureabstractSequential Mpeg-4 solutions actually developed for single processors try to integrate the most functionalities as possible in an unique software, and are generally oversized compared with the actual service requirement. Moreover, they can hardly be projected onto multiprocessors targets, leading to an extra load of source code and calculations, but also to a sub-optimal use of the architecture parallelism. This paper introduces a distributed Mpeg-4 application, where the system part is hosted by a standard PC, and the video decoder is supported by a multi-DSPs board. In particular, we present our AVSynDEx methodology allowing both an incremental building, an easy update on the video decoder description, and a quasi-automatic implementation onto a multi-C6x platform. We also define a global scheduler managing the parallel execution of the video and system applications. Nicolas Ventroux, Jean-François Nezan, Mickaël Raulet, Olivier Déforges |
ICME | 4 |
| 2002 | Supervised segmentation at low bit rates for region representation and color image compressionabstractWe present a new coding scheme providing compression efficiency (better subjective quality than JPEG2000) and region-level functionalities at low bit-rates for both the sender and the receiver. A contourless coding region description is obtained by a segmentation process at the sender and receiver, supervised by chromatic consistency and operating on a well adapted low resolution Y image. Chromatic images can then be region-based encoded at very low cost. As the region topology fully matches the representation used to encode the content, further enhancement can be globally applied over the whole image or limited to a region of interest. Olivier Déforges, Joseph Ronsin |
ICME (1) | 1 |
| 2002 | Rapid prototyping methodology for multi-DSP TI C6X platforms applied to an Mpeg-2 coding applicationabstractReal time signal and image applications have very important time constraints, involving the use of several powerful numerical calculation units. Our aim is to develop a fast prototyping process dedicated to parallel architectures made of several last generation Texas Instruments TMS320C6X DSP. The methodology is based on the use of SynDEx, a CAD software improving the algorithm implementation onto multiprocessor architectures, finding the best matching between an algorithm and an architecture. A SynDEx executive kernel has been developed for the C6X DSP family in order to automatically generate a distributed and optimized static executive of the specified algorithm onto those processors. We have tested the efficiency of our methodology with a complete Mpeg-2 coding application. Jean-François Nezan, Olivier Déforges, Mickaël Raulet |
SPAA | 2 |
| 2002 | ARIAL: Rapid Prototyping for Mixed and Parallel Platforms
Virginie Fresse, Olivier Déforges |
Parallel Comput. | 2 |
| 2000 | A generic systolic processor for real time grayscale morphologyabstractMathematical morphology is a powerful and widespread tool for image analysis and coding, but for large analysis neighborhoods, it becomes rapidly time consuming on general purpose machines. Much work has been done to decrease the overall complexity by splitting heavy computation tasks into several more efficient ones. Even if significant speedup factors have been achieved using dedicated architectures, the major drawback is that the number of operations required is still dependent on the considered neighborhood size. This paper introduces a systolic network for real time grey level morphological processing since it is able to perform erosions, dilations, openings, and closings during a unique image pass whatever the neighborhood shape and size. It is based on a new structuring elements recursive decomposition by two pixels dilations and union sets. The resulting architecture is fully regular, built from elementary interconnected stages. By pipelining these stages timing performances are independent of the structuring element complexity. Efficient synthesis has been achieved into a FPGA enabling dual port ram implementation. Olivier Déforges, Nicolas Normand |
ICASSP | 1 |
| 2000 | Rapid prototyping for mixed architecturesabstractThe aim of our work is to achieve a rapid prototyping dedicated to mixed architectures, made up of one multi-DSP part (software) and a FPGA architecture (dedicated hardware). We have determined a complete codesign methodology enabling to implement a complete digital signal or image processing line. The starting description associated with this methodology, is only a functional description, represented by a data flow graph. This global process developed for an automatic implementation enables the user to develop complex applications at a high level onto a complex architecture without any implementation pre-requirements. Virginie Fresse, Olivier Déforges, Mustapha Assouil |
ICASSP | 2 |
| 1997 | Recursive Morphological Operators for Gray Image Processing. Application in Granulometry AnalysisabstractThis paper presents a new algorithm for an efficient implementation of morphological operations for gray images. It defines a recursive morphological decomposition method of convex structuring elements by only causal two pixel structuring elements. Whatever the element size, erosion or/and dilation can then be performed during a unique raster-like image scan, involving a fixed reduced analysis neighborhood. The resulting process offers a low computational complexity, combined with an easiness for describing the element form. The algorithm is exemplified with granulometry. Quantum dots are segmented using a multiscale morphologic decomposition. Our new algorithm is particularly well suited for this type of morphological treatments, as they use structuring elements with both a large size and a form fitting the object to extract, that is to say depending on the application. Olivier Déforges, Nicolas Normand |
ICIP (2) | 1 |
| 1995 | A Design Tool for the Specification and the Simulation of Array Processors Architectures - Application to Image Processing: The Extraction of Regions of InterestsabstractThis paper deals with a CAD tool dedicated to the design and the simulation of specific array processor architectures. These architectures are described into a specific notation which includes major characteristics of the VHDL syntax. This language provides a very concise and legible means to specify array processors. A preprocessor generates full standard VHDL code describing the behavior of the designed architecture. An original application to image processing is given: the design of a specific architecture for the extraction of regions of interests. Gérard Ramstein, Olivier Déforges, Przemyslaw Bakowski |
ASAP | 2 |
| 1995 | Segmentation of Complex Documents Multilevel Images: A Robust and Fast Text Bodies-Headers Detection and Extraction Scheme 770abstractWe present a method for segmenting multilevels images of documents. The documents are considered difficult ones in the sense they may contain text paragraphs with different orientations and shapes, mixed with graphics and photographs. The proposed method extracts and separates blocks of text lines (printed or handwritten characters) and headers as well as stroke structures. The generic approach is first based on a multiscale analysis with the use of a pyramid representation of the image. At each level, text location is performed by a line borders detection scheme. Then, an efficient bottom-up procedure generates bodies (text paragraphs) as the output of algebric transformations upon a set of four directed graphs associated with the topological relationships of physical components. Olivier Déforges, Dominique Barba |
ICDAR | 1 |
| 1994 | A Fast Multiresolution Text-Line and Non Text-Line Structures ExtractionabstractThis work is devoted to the description and the evaluation of a generic method for locating text and lines in document images. One of the possible applications of this scheme is to detect and carefully extract the destination address block (DAB) on flat mail pieces. Concerning the low level design stages, the methodology is based on a multiresolution description of the document image and a joint extraction of both text lines and non text lines using their global aspect. Intermediate and high level stages consist of grouping text words into block structures, and this is performed in a efficient and robust way relying not only on some lines properties, but also on the adjacent lines.> Olivier Déforges, Dominique Barba |
ICIP (1) | 1 |
| 1994 | A robust and multiscale document image segmentation for block line/text line structures extractionabstractThis paper presents a new approach dealing with the extraction of blocks of lines and with discrimination between the text line and nontext lines structures. This approach combines a multiresolution textural description of the document image with a cooperation between adjacent resolution levels. At the same time, a not purely bottom-up structural procedure allows us to construct in a very robust way text lines and nontext lines structures and group them into fairly homogeneous blocks. The efficiency of the overall method is demonstrated by its ability to segment and extract destination address block (DAB) on quite diverse and complex images of flat mail pieces. Olivier Déforges, Dominique Barba |
ICPR (1) | 1 |