Frédéric Dufaux

dblp:15/3172 · DBLP profile ↗
← Back
135ranked-venue papers
14as first author
29since 2021 · last 2026
0000-0001-6388-4112ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 128 · 12 first-author · 28 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Systems, architecture and hardware · 2 · 1 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Symmetric Entropy-Constrained Video Coding for Machines
abstract
As video transmission increasingly serves machine vision systems (MVS) instead of human vision systems (HVS), video coding for machines (VCM) has become a critical research topic. Existing VCM methods often bind codecs to specific downstream models, requiring retraining or supervised data, thus limiting generalization in multi-task scenarios. Recently, unified VCM frameworks have employed visual backbones (VB) and visual foundation models (VFM) to support multiple video understanding tasks with a single codec. They mainly utilize VB/VFM to maintain semantic consistency or suppress non-semantic information, but seldom explore how to directly link video coding with understanding under VB/VFM guidance. Hence, we propose a Symmetric Entropy-Constrained Video Coding framework for Machines (SEC-VCM). It establishes a symmetric alignment between the video codec and VB, allowing the codec to leverage VB's representation capabilities to preserve semantics and discard MVS-irrelevant information. Specifically, a bi-directional entropy-constraint (BiEC) mechanism ensures symmetry between the process of video decoding and VB encoding by suppressing conditional entropy. This helps the codec to explicitly handle semantic information beneficial to MVS while squeezing useless information. Furthermore, a semantic-pixel dual-path fusion (SPDF) module injects pixel-level priors into the final reconstruction. Through semantic-pixel fusion, it suppresses artifacts harmful to MVS and improves machine-oriented reconstruction quality. Experimental results on classical video understanding tasks and MLLM-based tasks show state-of-the-art (SOTA) rate-task performance. It achieves significant bitrate savings over H.266/VVC reference software VTM on video instance segmentation (37.4%), video object segmentation (29.8%), object detection (46.2%), multiple object tracking (44.9%), and MLLM-based video grounding (97.6%). The code is at https://github.com/Ws-Syx/SEC-VCM.
Yuxiao Sun, Meiqin Liu 0002, Weisi Lin, Frédéric Dufaux, Yao Zhao 0001
IEEE Trans. Image Process.7
2025 Channel and space-based joint rate allocation algorithm
abstract
Rate control is a critical component for image and video compression Particularly under limited network bandwidth conditions, bitrate control is essential to ensure efficient image transmission by effectively allocation channel resources. In this research, since both Channel and Spatial have relationship with rate allocation, we first propose a joint Channel-wise and Spatial-wise Quantization scheme to determine optimal quantization parameters. Subsequently, we develop a quantization step estimation network to obtain parameters to efficiently allocate rate according to target rate. Experiments demonstrate that our algorithm significantly improve compressed image quality with minimal bitrate distortion and achieve accurate rate control with nearly 3% average bitrate error.
Yu Sun 0003, Xin Lu 0001, Frédéric Dufaux, Ce Zhu
ICASSP6
2025 Lift-PCAC: Lifting Based Point Cloud Attribute Compression
abstract
Point cloud (PC) compression is crucial for efficient transmission and storage in applications like virtual and augmented reality, where point counts can reach millions. While learning-based methods have shown promise in compressing PC geometry, attribute compression remains relatively unexplored. Existing methods often rely on variational autoencoders (VAEs). However, VAEs, with their low-dimensional bottlenecks, inherently limit the achievable reconstruction quality, especially at high bitrates. In this paper we introduce a novel approach to compressing PC attributes using a lifting framework. The invertibility of our method enables better reconstruction quality at high bitrates. Our Lifting Based Point Cloud Attribute Compression (Lift-PCAC) outperforms existing learning-based attribute compression methods on higher bitrates and shows comparable performance to G-PCC v.21 in some cases, highlighting the potential of this approach for point cloud compression.
Rodrigo B. Pinheiro, Jean-Eudes Marvie, Giuseppe Valenzise, Frédéric Dufaux
ICIP4
2025 Parallel-Based Fast Coding Mode Decision for Intra Coding in VVC SCC
abstract
In light of the growing popularity of screen content video applications, there is a increasing demand for Screen Content Coding (SCC). The latest standard, Versatile Video Coding (VVC), exhibits exceptionally high coding efficiency, albeit accompanied by a considerable coding complexity. This complexity, in turn, restricts the widespread applicability of VVC SCC. To address this issue, this paper introduces a parallel based approach to enhance the coding speed of VVC SCC Intra Coding. Specifically, we established a large-scale database and then design distinct neural networks for Coding Units (CUs) of various sizes to predict candidate Coding Modes (CMs). Subsequently, we formulate a loss function based on CM distributions and Rate Distortion(RD) costs to train the designed models. Finally, we introduce a threshold selection scheme to balance coding efficiency and coding speed. Experimental results demonstrate that the proposed method improves coding speed by an average of 36.36%, with an average increase of 0.95% in Bjøntegaard Delta Bit Rate (BDBR).
Kongqing Peng, Xin Lu 0001, Frédéric Dufaux, Shibin Zhang, Weian Li, Hongwei Guo 0001
ICIP4
2025 A Multi-Layer End-to-End 360 Image Compression
abstract
360° images have attracted increasing attention due to their wide field of view. However, spherical 360° images need to be converted into 2D equi-rectangular projection (ERP) images for compression. This conversion often leads to pixel overstretching in the ERP image, which results in a lot of redundancy in the texture. Performing a direct prediction without considering the stretching will inevitably make it difficult to achieve optimal results. To tackle this problem, we propose a multi-layer adaptive scale-block scheme for ERP image compression. In particular, we introduce a multi-layer structure based on the overstretching rate and use multi-scale convolution kernels to better match each layer and extract features more effectively. Subsequently, we employ an adaptive scale-block method to effectively reduce bitrate redundancy in overstretched and less important regions. Finally, we propose a new end-to-end model that is efficient for 360° image compression. Experimental results demonstrate that our scheme outperforms other image compression methods and reduces bitrate by nearly 16% compared to the latest learned 360° image compression model.
Yubiao Zhou, Yu Sun 0003, Frédéric Dufaux, Weian Li, Ce Zhu
ICIP4
2025 Fast CU Partition Algorithm For 360-Degree Videos on VVC
abstract
360-degree videos (abbreviated as 360 videos) have gained widespread popularity due to their immersive experience. The massive amount of video data resulting from ultra-high resolution of 360 videos makes the coding process extremely complex and seriously limits their widespread applications. In this paper, we propose a new fast-Coding Unit (CU) partition algorithm for 360 videos based on Versatile Video Coding (VVC). The major novelty is that we establish databases and develop models based on stretching and split mode (SM) distribution for 360 videos on VVC. Specifically, first, we establish stretching-based CU partition databases for 360 videos. Second, we propose a new stretching-based multi-scale convolution kernel network structure and a corresponding synthetical loss function, which further considers both Rate-Distortion cost and imbalance issue. We then develop a unique multi-threshold selection scheme to select candidate SMs. Experimental results demonstrate that the proposed algorithm improves encoding speed by 69.63%, with only a 2.28% increase in Bjøntegaard Delta Bit Rate (BDBR), outperforming the state-of-the-art methods.
Shijie Du, Yu Sun 0003, Shuyin Xia, Frédéric Dufaux, Hongwei Guo 0001, Guoyin Wang 0001, Ce Zhu
ICME5
2025 DA-Flow: Dual Attention Normalizing Flow for Skeleton-Based Video Anomaly Detection
abstract
Cooperation between temporal convolutional networks (TCN) and graph convolutional networks (GCN) as a processing module has shown promising results in skeleton-based video anomaly detection (SVAD). However, to maintain a lightweight model with low computational and storage complexity, shallow GCN and TCN blocks are constrained by small receptive fields and a lack of cross-dimension interaction capture. To tackle this limitation, we propose a lightweight module called the Dual Attention Module (DAM) for capturing cross-dimension interaction relationships in spatio-temporal skeletal data. It employs the frame attention mechanism to identify the most significant frames and the skeleton attention mechanism to capture broader relationships across fixed partitions with minimal parameters and total Floating Point Operations (FLOPs). Furthermore, the proposed Dual Attention Normalizing Flow (DA-Flow) integrates the DAM as a post-processing unit after GCN within the normalizing flow framework. Simulations show that the proposed model is robust against noise and negative samples. Experimental results show that DA-Flow reaches competitive or better performance than the existing state-of-the-art (SOTA) methods in terms of the micro AUC metric with the fewest parameters and FLOPs. Moreover, we found that even without training, simply using random projection without dimensionality reduction on skeleton data enables substantial anomaly detection capabilities.
Ruituo Wu, Bing Li 0002, Jicong Fan 0001, Frédéric Dufaux, Ce Zhu, Yipeng Liu 0001
IEEE Trans. Multim.6
2024 Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics
abstract
Automatic metrics are used as proxies to evaluate abstractive summarization systems when human annotations are too expensive.To be useful, these metrics should be fine-grained, show a high correlation with human annotations, and ideally be independent of reference quality; however, most standard evaluation metrics for summarization are reference-based, and existing reference-free metrics correlate poorly with relevance, especially on summaries of longer documents.In this paper, we introduce a reference-free metric that correlates well with human evaluated relevance, while being very cheap to compute.We show that this metric can also be used alongside reference-based metrics to improve their robustness in low quality reference settings.
Théo Gigant, Camille Guinaudeau, Marc Décombas, Frédéric Dufaux
EMNLP4
2024 Reducing the Complexity of Normalizing Flow Architectures for Point Cloud Attribute Compression
abstract
Existing learning-based methods to compress PCs attributes typically employ variational autoencoders (VAE) to learn compact signal representations. However, these schemes suffer from limited reconstruction quality at high bitrates due to their intrinsic lossy nature. More recently, normalizing flows (NF) have been proposed as an alternative solution. NFs are invertible networks that can achieve lossless reconstruction, at the cost of very large architectures with high memory and computational footprint. This paper proposes an improved NF architecture with reduced complexity called RNF-PCAC. It is composed of two operating modes specialized for low and high bitrates, combined in a rate-distortion optimized fashion. Our approach reduces the number of parameters of the existing NF architectures by over 6×. At the same time, it achieves state-of-the-art coding gains compared to previous learning-based methods and, for some PCs, it matches the performance of G-PCC (v.21).
Rodrigo B. Pinheiro, Jean-Eudes Marvie, Giuseppe Valenzise, Frédéric Dufaux
ICASSP4
2024 Balancing Representation Abstractions and Local Details Preservation for 3d Point Cloud Quality Assessment
abstract
3D Point Clouds (PCs) have become a valuable tool for representing intricate 3D information. Assessing the quality of PCs remains a challenging task, especially when striving for optimal immersive experiences. This paper introduces a novel metric and training approach that leverages projection-based views to evaluate the quality of 3D content. Our approach addresses a critical issue related to the intrinsic bias of deep networks for image recognition towards building hierarchical representations including only the global semantic, at the expense of local details. This bias is a limiting factor in tasks like 3D point cloud quality assessment where instances of the same content with varying degrees and types of degradation can possess strikingly similar representations. We propose a novel point cloud quality metric using a dual supervised and unsupervised training strategy to balance semantic understanding and preservation of critical perceptual quality-relevant information. The results demonstrate the effectiveness and reliability of our solution compared to state-of-the-art metrics on two standard 3D PCs quality assessment benchmarks (3D PCQA).
Marouane Tliba, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux
ICASSP4
2024 Fast Intra Mode Prediction Algorithms for SCBS in VVC SCC
abstract
Versatile Video Coding (VVC) now supports Screen Content Coding (SCC) by integrating two efficient coding modes: Intra Block Copy (IBC) and Palette (PLT). However, the numerous modes and the Quad-Tree Plus Multi-Type Tree (QTMT) structure inherent to VVC contribute to a very high coding complexity. To effectively reduce the computational complexity of VVC SCC, we propose a fast Intra mode prediction algorithm for VVC SCC. More specifically, we first use the difference of minimum Sum of Absolute Transformed Differences (SATD) value of four Directional Modes (DMs) of Intra and the SATD value of the IBC-merge mode to determine whether to early skip Intra checking. Subsequently, we use a decision tree to determine whether to early terminate the checking after block differential pulse coded modulation (BDPCM). Finally, we employ a decision tree to determine whether to early skip multiple transform selection (MTS) and low frequency non-separable transform (LFNST) checking. The results demonstrate that our algorithm achieves an average encoding time reduction of 34.34% with a negligible Bjøntegaard delta bitrate increase of 0.46%.
Yishen Deng, Weisheng Li 0001, Xin Lu 0001, Frédéric Dufaux, Bo Hang, Ce Zhu
ICASSP5
2024 Fast Coding Mode Prediction for Intra Prediction in VVC SCC
abstract
Currently, screen content video applications are increasingly widespread in our daily lives. The latest Screen Content Coding (SCC) standard, known as Versatile Video Coding (VVC) SCC, employs screen content Coding Modes (CMs) selection. While VVC SCC achieves high coding efficiency, its coding complexity poses a significant obstacle to the further widespread adoption of screen content video. Hence, it is crucial to enhance the coding speed of VVC SCC. In this paper, we propose a fast mode and splitting decision for Intra prediction in VVC SCC. Specifically, we initially exploit deep learning techniques to predict content types for all CUs. Subsequently, we examine CM distributions of different content types to predict candidate CMs for CUs. We then introduce early skip and early terminate CM decisions for different content types of CUs to further eliminate unlikely CMs. Finally, we develop Block-based Differential Pulse-Code Modulation (BDPCM) early termination to improve coding speed. Experimental results demonstrate that the proposed algorithm can improve coding speed by $34.95 \%$ on average while maintaining almost the same coding efficiency.
Junyi Yu, Xin Lu 0001, Frédéric Dufaux, Hongwei Guo 0001, Ce Zhu
ICIP4
2024 Enhancing Immersive Experiences through 3D Point Cloud Analysis: A Novel Framework for Applying 2D Visual Saliency Models to 3D Point Clouds
abstract
In the new area of immersive multimedia environments, understanding and manipulating visual attention are crucial for enhancing user experience. This study introduces an innovative framework that extends traditional 2D saliency maps to the analysis of 3D point clouds, a step forward in adapting saliency prediction to more complex and immersive environments. Our framework centers on the orthographic projection of 3D point clouds onto 2D planes, enabling the application of established 2D saliency models to this novel context. We further delve into the evaluation of these models on a 3D point cloud eye-tracking dataset, exploring various projection settings and thresholding techniques to maintain the integrity of saliency information in the transition from 2D to 3D. This research not only bridges a gap in applying visual attention models to 3D data but also offers insights into the optimization of quality of experience in immersive multimedia systems.
Marouane Tliba, Xuemei Zhou, Irene Viola 0001, Pablo César, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux
QoMEX7
2023 TIB: A Dataset for Abstractive Summarization of Long Multimodal Videoconference Records
abstract
Large language models and multimodal language-vision models give impressive results on current available summarization benchmarks, but are not designed to handle long multimodal documents. Most summarization datasets are composed of either mono-modal documents or short multimodal documents. In order to develop models designed for understanding and summarizing real-world videoconference records that are typically around 1 hour long, we propose a dataset of 9,103 videoconference records extracted from the German National Library of Science and Technology (TIB) archive, along with their abstract. Additionally, we process the content using automatic tools in order to provide the transcripts and key frames. Finally, we present experiments for abstractive summarization, to serve as baseline for future research work in multimodal approaches.
Théo Gigant, Frédéric Dufaux, Camille Guinaudeau, Marc Décombas
CBMI2
2023 NF-PCAC: Normalizing Flow Based Point Cloud Attribute Compression
abstract
Learning-based point cloud (PC) compression is a promising research avenue to reduce the transmission and storage costs for PC applications. Existing learning-based methods to compress PCs have mainly focused on geometry and employ variational autoencoders to learn compact signal representations. However, autoencoders leverage low-dimensional bottlenecks that limit the maximum reconstruction quality, even at high bitrates. In this paper, we propose a different and novel approach to compress PC attributes by using normalizing flows. Since normalizing flows model invertible transforms, the proposed approach can achieve better reconstruction quality than variational autoencoders over a large range of bitrates. Our Normalizing Flow-based Point Cloud Attribute Compression (NF-PCAC) outperforms previous learning-based methods for attribute compression, and has comparable performance as G-PCC v.14, showing the potential of this scheme for PC compression.
Rodrigo B. Pinheiro, Jean-Eudes Marvie, Giuseppe Valenzise, Frédéric Dufaux
ICASSP4
2023 PCQA-Graphpoint: Efficient Deep-Based Graph Metric for Point Cloud Quality Assessment
abstract
Following the advent of immersive technologies and the increasing interest in representing interactive geometrical format, 3D Point Clouds (PC) have emerged as a promising solution and effective means to display 3D visual information. In addition to other challenges in immersive applications, objective and subjective quality assessments of compressed 3D content remain open problems and an area of research interest. Yet most of the efforts in the research area ignore the local geometrical structures between points representation. In this paper, we overcome this limitation by introducing a novel and efficient objective metric for Point Clouds Quality Assessment, by learning local intrinsic dependencies using Graph Neural Network (GNN). To evaluate the performance of our method, two well-known datasets have been used. The results demonstrate the effectiveness and reliability of our solution compared to state-of-the-art metrics.
Marouane Tliba, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux
ICASSP4
2023 A Novel Mode Selection-Based Fast Intra Prediction Algorithm for Spatial SHVC
abstract
Due to multi-layer encoding and Inter-layer prediction, Spatial Scalable High-Efficiency Video Coding (SSHVC) has extremely high coding complexity. It is very crucial to improve its coding speed so as to promote widespread and cost-effective SSHVC applications. In this paper, we have proposed a novel Mode Selection-Based Fast Intra Prediction algorithm for SSHVC. We reveal the RD costs of Inter-layer Reference (ILR) mode and Intra mode have a significant difference, and the RD costs of these two modes follow Gaussian distribution. Based on this observation, we propose to apply the classic Gaussian Mixture Model and Expectation Maximization in machine learning to determine whether ILR is the best mode so as to skip the Intra mode. Experimental results demonstrate that the proposed algorithm can significantly improve the coding speed with negligible coding efficiency loss.
Yu Sun 0003, Weisheng Li 0001, Lele Xie, Xin Lu 0001, Frédéric Dufaux, Ce Zhu
ICASSP6
2023 Fast Learning-Based Split Type Prediction Algorithm for VVC
abstract
As the latest video coding standard, Versatile Video Coding (VVC) is highly efficient at the cost of very high coding complexity, which seriously hinders its widespread application. Therefore, it is very crucial to improve its coding speed. In this paper, we propose a learning-based fast split type (ST) prediction algorithm for VVC using a deep learning approach. We first construct a large-scale database containing sufficient STs with diverse video resolution and content. Next, since the ST distributions of coding units (CUs) of different sizes are significantly distinct, so we separately design neural networks for all different CU sizes. Then, we merge ambiguous STs into four merged classes (MCs) to train models to obtain probabilities of MCs and skip unlikely ones. Experimental results demonstrate that the proposed algorithm can reduce the encoding time of VVC by 67.53% with 1.89% increase in Bjøntegaard delta bit-rate (BDBR) on average.
Liulin Chen, Xin Lu 0001, Frédéric Dufaux, Weisheng Li 0001, Ce Zhu
ICIP4
2023 A Probability-Based All-Zero Block Early Termination Algorithm for QSHVC
abstract
To seamlessly adapt to time-varying network bandwidths, Quality Scalable High-Efficiency Video Coding (QSHVC) is developed. However, its coding process is overwhelmingly complex, and this seriously limits its wide applications in realtime environments. Therefore, it is of great significance to study fast coding algorithms for QSHVC. In this paper, we propose a novel probability-based All-Zero Block (AZB) early termination algorithm for QSHVC. We observe that the generated residual coefficients follow the Laplace distribution if a CU is accurately predicted. Based on this observation, we derive the sum of squared differences-based AZB decision condition. Second, the probability of each coding mode and coding depth being chosen as the best ones are combined with AZBs to derive the probability-based early termination condition. The experimental results show that the proposed algorithm can improve the average coding speed by 74.95% with a 0.26% increase in BDBR.
Xin Lu 0001, Frédéric Dufaux, Qianmin Wang, Weisheng Li 0001, Bo Hang, Ce Zhu
ICIP3
2023 RV-TMO: Large-Scale Dataset for Subjective Quality Assessment of Tone Mapped Images
abstract
Tone mapping operators (TMO) are functions that map high dynamic range (HDR) images to a standard dynamic range (SDR), while aiming to preserve the perceptual cues of a scene that govern its visual quality. Despite the increasing number of studies on quality assessment of tone mapped images, current subjective quality datasets have relatively small numbers of images and subjective opinions. Moreover, existing challenges in transferring laboratory experiments to crowdsourcing platforms put a barrier for collecting large-scale datasets through crowdsourcing. In this work, we address these challenges and propose the RealVision-TMO (RV-TMO), a large-scale tone mapped image quality dataset. RV-TMO contains 250 unique HDR images, their tone mapped versions obtained using four TMOs and pairwise comparison results from seventy unique observers for each pair. To the best of our knowledge, this is the largest dataset available in the literature for quality evaluation of TMOs by the number of tone mapped images and number of annotations. Furthermore, we provide a content selection strategy to identify interesting and challenging HDR images. We also propose a novel methodology for observer screening in pairwise experiments. Our work does not only provide annotated data to benchmark existing objective quality metrics, but also paves the path to building new metrics for tone mapping quality evaluation.
Ali Ak, Abhishek Goswami, Wolf Hauser, Patrick Le Callet, Frédéric Dufaux
IEEE Trans. Multim.5
2022 A Multiscale Gradient-Backpropagation Optimization Framework for Deformable Convolution Based Compressed Video Enhancement
abstract
Deep learning based compressed video quality enhancement has raised lots of interest recently. To explore the information over multiple frames, deformable convolution has been used for temporal alignment. However, in the existing methods, the deformable convolution is used in a relatively naïve way, without differing the characteristics of offset and features, and their behavior in gradient backpropagation. In this paper, a multiscale gradient-backpropagation optimization framework is proposed for the deformable convolution based compressed video quality enhancement. By analyzing the gradient backpropagation mechanism of deformable convolution, a multi-scale deformable convolution alignment structure is developed to facilitate the gradient backpropagation at all scales. Moreover, a progressive offset prediction module is developed, which decouples the offset prediction from the feature up-sampling, thus reducing the noise flow over scales. Experimental results show that the proposed method achieves the state-of-the-art performance, with 25.6% BD-rate saving compared to the HEVC reference software (HM).
Yanbo Gao, Menghu Jia, Shuai Li 0005, Mao Ye 0001, Frédéric Dufaux
ICASSP6
2022 Representation Learning Optimization for 3D Point Cloud Quality Assessment Without Reference
abstract
Recent information and communication systems have employed 3D Point Cloud (PC) as an advanced geometrical representation modality for immersive applications. Like most multimedia data, PCs are often compressed for transmission and viewing purposes, which can impact the perceived quality. Developing robust and efficient objective quality metrics for PCs is still an open problem. In this paper, we propose an end-to-end deep approach for evaluating the perceptual effects of point cloud compression solutions without reference. Our approach focuses on leveraging the intrinsic point cloud characteristics to quantify the coding impairments from few distant randomly selected patches using supervised and unsupervised training strategies. To evaluate the performance of our method, two well-known datasets have been used. The results demonstrate the effectiveness and reliability of the proposed method compared to to state-of-the-art methods.
Marouane Tliba, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux
ICIP4
2022 Gaussian Distribution-based Mode Selection for Intra Prediction of Spatial SHVC
abstract
Due to the diversity of terminal devices, Spatial Scalable High Efficiency Video Coding (SSHVC) is an efficient solution to meet this requirement. However, its coding process is very complex, which seriously prevents its wide applications. Therefore, it is very crucial to reduce coding complexity and improve coding speed. In this paper, we propose a Gaussian Distribution-based Mode Selection for Intra Prediction of SSHVC. We show that the rate distortion costs of Inter-layer Reference (ILR) mode and Intra mode are significantly different, and both follow a Gaussian distribution. Based on this discovery, we propose to use a Bayes decision rule to determine whether ILR is the best mode so as to skip Intra mode. Experimental results demonstrate that the proposed algorithm can significantly improve coding speed with negligible coding efficiency losses.
Yu Sun 0003, Weisheng Li 0001, Xin Lu 0001, Frédéric Dufaux
ICIP6
2021 Interframe-Dependent Rate-QP-Distortion Model For Video Coding And Transmission
abstract
In this paper, we propose a new inter-dependent Rate-QP-Distortion model. This model predicts the size of the picture after compression based on the distortion (D) of the reference frame and the current Quantization parameter (QP). This model is particularly useful when adjusting the QP of the picture according to the allocated bitrate budget. Simulation results demonstrate that the proposed model outperforms other models in the literature. In the video sequence Tango, up to 90% of all prediction errors are inferior to 8.6% when using constant QP encoding, and 90% of all prediction errors are inferior to 12% when using variable QP encoding. One application of this model is in low latency video streaming, where each frame of the video sequence needs to be coded with a specific target bitrate, due to variations of the instantaneous transmission rate.
Mourad Aklouf, Marc Leny, Michel Kieffer, Frédéric Dufaux
ICIP4
2021 Convolutional Neural Network for 3D Point Cloud Quality Assessment with Reference
abstract
In recent years, the production of 3D content in the form of point clouds (PC) has increased considerably, especially in virtual reality applications. This enthusiasm is linked in particular to the development of acquisition technologies. In order to ensure a good quality of user experience, it is necessary to offer a high quality of visualization whatever the transmission medium used or the treatments applied. Thus, several metrics have been proposed which are essentially point-based metrics. In this article, we propose a deep learning-based method that efficiently predicts the quality of distorted PCs thanks to a set of features extracted from selected patches of the reference PC and its degraded version as well as the use of Convolutional Neural Networks (CNNs). The patches are selected randomly and the difference between corresponding patches is characterized by three attributes: geometry, curvature and color. The proposed method was evaluated and compared to state-of-the-art metrics using two datasets, including a large dataset more suited to deep learning models. We also compared different symmetrization functions and machine learning pooling as well as the ability of our method to predict the quality of unknown PCs through a cross-dataset evaluation. The results obtained show the relevance of the proposed framework with interesting perspectives.
Aladine Chetouani, Maurice Quach, Giuseppe Valenzise, Frédéric Dufaux
MMSP4
2021 Reliability of Crowdsourcing for Subjective Quality Evaluation of Tone Mapping Operators
abstract
Tone mapping operators (TMO) are functions which map high dynamic range (HDR) images to limited dynamic media while aiming to preserve the perceptual cues of the scene that govern its aesthetic quality. Evaluating aesthetic quality of TMOs is non-trivial due to the high subjectivity of preference involved. Traditionally, TMO aesthetic quality has been evaluated via subjective experiments in a controlled laboratory environment. However, the last decade has brought a surge in popularity of crowdsourcing as an alternative methodology to conduct subjective experiments. However, uncontrolled experiment conditions and unreliability of participant behaviour puts doubts on the trustworthiness of the collected data. In this study, we explore the possibility of using crowdsourcing platforms for subjective quality evaluation of TMOs. We have conducted three experiments with systematic changes to investigate the effect of experiment conditions and participant recruitment methods on the collected subjective data. Our results show that subjective evaluation of TMO aesthetic quality can be conducted on Prolific crowdsourcing platform with negligible differences in comparison to laboratory experiments. Furthermore, we provide objective conclusions about the effect of number of observers on the certainty of the pairwise comparison results.
Abhishek Goswami, Ali Ak, Wolf Hauser, Patrick Le Callet, Frédéric Dufaux
MMSP5
2021 Learning-based lossless light field compression
abstract
We propose a learning-based method for lossless light field compression. The approach consists of two steps: first, the view to be compressed is synthesized based on previously decoded views; then, the synthesized view is used as a context to predict probabilities of the residual signal for adaptive arithmetic coding. We leverage recent advances in deep-learning-based view synthesis and generative modeling. Specifically, we evaluate two strategies for entropy modeling: a fully parallel probability estimation, where all pixel probabilities are estimated simultaneously; and a partially auto-regressive estimation, in which groups of pixels are predicted sequentially. Our results show that the latter approach provides the best coding gains compared to the state of the art, while keeping the computational complexity competitive.
Milan Stepanov, M. Umair Mukati, Giuseppe Valenzise, Søren Forchhammer, Frédéric Dufaux
MMSP5
2021 Query-by-example HDR image retrieval based on CNN
Raoua Khwildi, Azza Ouled Zaid, Frédéric Dufaux
Multim. Tools Appl.3
2021 Guest Editorial Introduction to the Special Issue on Recent Advances in Point Cloud Processing and Compression
abstract
A point cloud is a set of 3D points that can be used to represent a 3D surface. Each point has a spatial position (x, y, z) and a vector of attributes, such as colors, material reflection, or normal. As point clouds are capable of reconstructing 3D objects or scenes, they have the potential to be widely used in various applications such as auto-driving and 6-degree virtual reality. However, the following properties of point cloud make the point cloud compression and processing become rather challenging. 1) Unstructured. The point cloud is a series of non-uniform sampled points. On the one hand, it makes the correlations among various points difficult to be utilized for compression. On the other hand, the convolutional neural network that is widely used in image/video processing cannot be applied to the point cloud processing. 2) Unordered. Unlike images and videos, the point cloud is a set of points without a specific order. Therefore, both the point cloud processing and compression algorithms need to be invariant to any permutations of the input point clouds.
Zhu Li 0001, Shan Liu 0001, Frédéric Dufaux, Li Li 0040, Ge Li 0002, C.-C. Jay Kuo
IEEE Trans. Circuits Syst. Video Technol.3
2020 Folding-Based Compression Of Point Cloud Attributes
abstract
Existing techniques to compress point cloud attributes leverage either geometric or video-based compression tools. We explore a radically different approach inspired by recent advances in point cloud representation learning. Point clouds can be interpreted as 2D manifolds in 3D space. Specifically, we fold a 2D grid onto a point cloud and we map attributes from the point cloud onto the folded 2D grid using a novel optimized mapping method. This mapping results in an image, which opens a way to apply existing image processing techniques on point cloud attributes. However, as this mapping process is lossy in nature, we propose several strategies to refine it so that attributes can be mapped to the 2D grid with minimal distortion. Moreover, this approach can be flexibly applied to point cloud patches in order to better adapt to local geometric complexity. In this work, we consider point cloud attribute compression; thus, we compress this image with a conventional 2D image codec. Our preliminary results show that the proposed folding-based coding scheme can already reach performance similar to the latest MPEG Geometry-based PCC (G-PCC) codec.
Maurice Quach, Giuseppe Valenzise, Frédéric Dufaux
ICIP3
2020 Just Noticeable Quantization Levels For High Dynamic Range Images
abstract
Just noticeable quantization levels, which are conventionally used in picture coding, have been mainly developed for standard 8-bit images and low dynamic range (LDR) typical screens. The quantization levels however have not been adapted yet for high dynamic range (HDR) imaging and its accompanied HDR displays, which can reach up to a peak luminance of 4000 cd/m2. This study proposes an experimental methodology on HDR displays to determine just noticeable quantization levels for discrete cosine transform (DCT) coefficients on high luminance images. In the first stage of the proposed method, the quantization noise patterns for different DCT frequencies at different mean luminances are rendered by predicting the LED and LCD values of the two layer HDR display. Then, a two alternative forced choice based psychovisual experimental procedure using geometric search and QUEST methodology is realized by randomly presenting the rendered quantization noise at different amplitudes to the subjects in order to determine the just noticeable levels. The experiments are performed over 3 subjects for 30 different frequencies of 8×8 DCT patterns at mean luminances of 100 cd/m2and 1000 cd/m2. The results are interpreted with respect to frequency and luminance changes and from the point of utilized methodology, namely geometric search and QUEST.
Sevim Begüm Sözer, Alper Koz, Ahmet Oguz Akyüz, Emin Zerman, Giuseppe Valenzise, Frédéric Dufaux
ICIP6
2020 Hybrid Learning-Based And Hevc-Based Coding Of Light Fields
abstract
Light fields have additional storage requirements compared to conventional image and video signals, and demand therefore an efficient representation. In order to improve coding efficiency, in this work we propose a hybrid coding scheme which combines a learning-based compression approach with a traditional video coding scheme. Their integration offers great gains at low/mid bitrates thanks to the efficient representation of the learning-based approach and is competitive at high bitrates compared to standard tools thanks to the encoding of the residual signal. The proposed approach achieves on average 38% and 31% BD rate saving compared to HEVC and JPEG Pleno transform-based codec, respectively.
Milan Stepanov, Giuseppe Valenzise, Frédéric Dufaux
ICIP3
2020 Improved Deep Point Cloud Geometry Compression
abstract
Point clouds have been recognized as a crucial data structure for 3D content and are essential in a number of applications such as virtual and mixed reality, autonomous driving, cultural heritage, etc. In this paper, we propose a set of contributions to improve deep point cloud compression, i.e.: using a scale hyperprior model for entropy coding; employing deeper transforms; a different balancing weight in the focal loss; optimal thresholding for decoding; and sequential model training. In addition, we present an extensive ablation study on the impact of each of these factors, in order to provide a better understanding about why they improve RD performance. An optimal combination of the proposed improvements achieves BD-PSNR gains over G-PCC trisoup and octree of 5.50 (6.48) dB and 6.84 (5.95) dB, respectively, when using the point-to-point (point-to-plane) metric. Code is available at https://github.com/mauriceqch/pcc_geo_cnn_v2.
Maurice Quach, Giuseppe Valenzise, Frédéric Dufaux
MMSP3
2020 Counter-examples generation from a positive unlabeled image dataset
Florent Chiaroni, Ghazaleh Khodabandelou, Mohamed-Cherif Rahal, Nicolas Hueber, Frédéric Dufaux
Pattern Recognit.5
2020 Deep Tone Mapping Operator for High Dynamic Range Images
abstract
A computationally fast tone mapping operator (TMO) that can quickly adapt to a wide spectrum of high dynamic range (HDR) content is quintessential for visualization on varied low dynamic range (LDR) output devices such as movie screens or standard displays. Existing TMOs can successfully tone-map only a limited number of HDR content and require an extensive parameter tuning to yield the best subjective-quality tone-mapped output. In this paper, we address this problem by proposing a fast, parameter-free and scene-adaptable deep tone mapping operator (DeepTMO) that yields a high-resolution and high-subjective quality tone mapped output. Based on conditional generative adversarial network (cGAN), DeepTMO not only learns to adapt to vast scenic-content (e.g., outdoor, indoor, human, structures, etc.) but also tackles the HDR related scene-specific challenges such as contrast and brightness, while preserving the fine-grained details. We explore 4 possible combinations of Generator-Discriminator architectural designs to specifically address some prominent issues in HDR related deep-learning frameworks like blurring, tiling patterns and saturation artifacts. By exploring different influences of scales, loss-functions and normalization layers under a cGAN setting, we conclude with adopting a multi-scale model for our task. To further leverage on the large-scale availability of unlabeled HDR data, we train our network by generating targets using an objective HDR quality metric, namely Tone Mapping Image Quality Index (TMQI). We demonstrate results both quantitatively and qualitatively, and showcase that our DeepTMO generates high-resolution, high-quality output images over a large spectrum of real-world scenes. Finally, we evaluate the perceived quality of our results by conducting a pair-wise subjective study which confirms the versatility of our method.
Aakanksha Rana, Praveer Singh, Giuseppe Valenzise, Frédéric Dufaux, Nikos Komodakis, Aljoscha Smolic
IEEE Trans. Image Process.4
2020 Fast Depth and Mode Decision in Intra Prediction for Quality SHVC
abstract
Scalable High Efficiency Video Coding (SHVC) is the extension of High Efficiency Video Coding (HEVC). In intra prediction for quality SHVC, a Coding Unit (CU) is recursively divided into a quadtree-based structure from the largest 64×64 CU to the smallest 8×8 CU, in which 35 intra prediction modes and Inter-Layer Reference (ILR) mode are checked to determine the best possible mode. This leads to very high coding efficiency but also results in an extremely high coding complexity. To improve coding speed while maintaining coding efficiency, in this paper, we propose a new efficient algorithm for fast intra prediction for enhancement layer in SHVC. First, temporal and spatial correlations, as well as their correlation degrees, are combined in a Naive Bayes classifier to predict depth probabilities and skip depths with low likelihood. Second, for a given depth candidate, we combine ILR mode probability with Partial Zero Blocks (PZBs) based on the Sum of Squared Differences (SSD) to determine whether the ILR mode is the best one. In that case, we can skip intra prediction, which requires very high complexity. Third, initial Intra Modes (IMs) are obtained through Sobel operator, and are combined with the relationship between IMs and their corresponding Hadamard Cost (HC) values to predict candidate IMs in Rough Mode Decision (RMD). Then, an analytical criterion of early termination is developed based on the HC values of two neighboring IMs in the Rate-Distortion Optimization (RDO) process. Finally, we combine depth probabilities and the distribution of residual coefficients at the current depth to early terminate depth selection. The proposed scheme can significantly decrease the complexity of depth determination while reducing the complexity of mode decision for a depth candidate. Our experimental results demonstrate that the proposed scheme can achieve a speed up gain of more than 80% in average, while maintaining coding efficiency.
Yu Sun 0003, Ce Zhu, Weisheng Li 0001, Frédéric Dufaux, Jiangtao Luo
IEEE Trans. Image Process.5
2020 Fast Depth and Inter Mode Prediction for Quality Scalable High Efficiency Video Coding
abstract
The scalable high efficiency video coding (SHVC) is an extension of high efficiency video coding (HEVC). It introduces multiple layers and inter-layer prediction, thus significantly increases the coding complexity on top of the already complicated HEVC encoder. In inter prediction for quality SHVC, in order to determine the best possible mode at each depth level, a coding tree unit can be recursively split into four depth levels, including merge mode, inter2N×2N, inter2N×N, interN×2N, interN×N, inter2N×nU, inter2N×nD, internL×2N and internRx×2N, intra modes and inter-layer reference (ILR) mode. This can obtain the highest coding efficiency, but also result in very high coding complexity. Therefore, it is crucial to improve coding speed while maintaining coding efficiency. In this research, we have proposed a new depth level and inter mode prediction algorithm for quality SHVC. First, the depth level candidates are predicted based on inter-layer correlation, spatial correlation and its correlation degree. Second, for a given depth candidate, we divide mode prediction into square and non-square mode predictions respectively. Third, in the square mode prediction, ILR and merge modes are predicted according to depth correlation, and early terminated whether residual distribution follows a Gaussian distribution. Moreover, ILR mode, merge mode and inter2N×2N are early terminated based on significant differences in Rate Distortion (RD) costs. Fourth, if the early termination condition cannot be satisfied, non-square modes are further predicted based on significant differences in expected values of residual coefficients. Finally, inter-layer and spatial correlations are combined with residual distribution to examine whether to early terminate depth selection. Experimental results have demonstrated that, on average, the proposed algorithm can achieve a time saving of 71.14%, with a bit rate increase of 1.27%.
Yu Sun 0003, Ce Zhu, Weisheng Li 0001, Frédéric Dufaux
IEEE Trans. Multim.5
2019 Transform Coefficient Coding for Screen Content in Versatile Video Coding (VVC)
abstract
A transform coefficient coding scheme is proposed for 4 × 4 blocks in Versatile Video Coding (VVC), targeting screen content applications. The proposed algorithm, called Unary Bitplane Coding (UBC), uses unary codes of the coefficient amplitudes and represents each block by their bitplanes. This representation allows exploiting further contextual information for source separation during the entropy coding. Experiments in the Joint Exploration test Model (JEM) show that replacing the existing transform coding with UBC only for 4 × 4 blocks brings on average 2.8% and 3.4% BD-R gain in the random access and all intra modes, respectively.
Mohsen Abdoli, Félix Henry, Patrice Brault, Frédéric Dufaux, Pierre Duhamel
ICASSP4
2019 Intra Block-DPCM with Layer Separation of Screen Content in VVC
abstract
An intra coding algorithm with layer separation is proposed. This algorithm is designed on top of an adopted tool in VVC, called Block DPCM (BDPCM), and benefits from texture information in a neighborhood to derive intensity levels of background and foreground layers. This information is used to reduce large rate of residual in case of incorrect layer prediction by BDPCM. For this purpose, three inter-layer transition states are defined that are either implicitly or explicitly conveyed to the decoder. Once a transition is signaled, the decoder corrects the prediction value using the derived layer information. Experiments on screen contents show a BD-rate gain of about 10% percent over VVC Test Model (VTM) and 1% over the regular BDPCM, with the cost of computational complexity.
Mohsen Abdoli, Félix Henry, Patrice Brault, Frédéric Dufaux, Pierre Duhamel, Pierrick Philippe
ICIP4
2019 Hallucinating A Cleanly Labeled Augmented Dataset from A Noisy Labeled Dataset Using GAN
abstract
Noisy labeled learning methods deal with training datasets containing corrupted labels. However, prediction performances of existing methods on small datasets still leave room for improvements. With this objective, in this paper we present a GAN-based method to generate a clean augmented training dataset from a small and noisy labeled dataset. The proposed approach combines noisy labeled learning principles with GAN state-of-the-art techniques. We demonstrate the usefulness of the proposed approach through an empirical study on simple and complex image datasets.
Florent Chiaroni, Mohamed-Cherif Rahal, Nicolas Hueber, Frédéric Dufaux
ICIP4
2019 Learning Convolutional Transforms for Lossy Point Cloud Geometry Compression
abstract
Efficient point cloud compression is fundamental to enable the deployment of virtual and mixed reality applications, since the number of points to code can range in the order of millions. In this paper, we present a novel data-driven geometry compression method for static point clouds based on learned convolutional transforms and uniform quantization. We perform joint optimization of both rate and distortion using a trade-off parameter. In addition, we cast the decoding process as a binary classification of the point cloud occupancy map. Our method outperforms the MPEG reference solution in terms of rate-distortion on the Microsoft Voxelized Upper Bodies dataset with 51.5% BDBR savings on average. Moreover, while octree-based methods face exponential diminution of the number of points at low bitrates, our method still produces high resolution outputs even at low bitrates. Code and supplementary material are available at https://github.com/mauriceqch/pcc_geo_cnn.
Maurice Quach, Giuseppe Valenzise, Frédéric Dufaux
ICIP3
2019 Fast Inter Mode Predictions for SHVC
abstract
The Scalable High Efficiency Video Coding (SHVC) has very high coding efficiency, but its computational complexity is also very high. This definitely limits its wide applications, particularly for real-time video applications. Therefore, it is crucial to improve the coding speed. In this research, we have proposed a new inter mode prediction algorithm for quality SHVC, in order to improve the coding speed while maintaining coding efficiency. First, we divide mode prediction into square mode prediction and non-square mode prediction. Second, in the square mode prediction, Inter-Layer Reference (ILR) and merge modes are predicted based on depth correlation. Moreover, ILR mode, merge mode and inter 2N×2N are early terminated based on Rate Distortion (RD) cost. Third, if the early termination condition cannot be satisfied, nonsquare modes are further predicted based on the distribution of residual coef-ficients. Experimental results have demonstrated that the proposed algorithm can significantly improve the coding speed with negligible coding efficiency loss.
Yu Sun 0003, Weisheng Li 0001, Ce Zhu, Frédéric Dufaux
ICME5
2019 Predicting Subjectivity in Image Aesthetics Assessment
abstract
Conventional image aesthetic quality prediction aims at predicting the average score of a picture or its aesthetic class (good/bad quality). However, aesthetic prediction is intrinsically subjective, and images with similar mean aesthetic scores/class might display very different levels of consensus by human raters. Recent work has dealt with aesthetic subjectivity by predicting the distribution of human scores. However, predicting the distribution is not directly interpretable in terms of subjectivity, and might be sub-optimal compared to directly estimating subjectivity descriptors computed from ground-truth scores. In this paper, we propose several measures of subjectivity, ranging from simple statistical measures such as the standard deviation of the scores, to newly proposed descriptors inspired by information theory. We evaluate the prediction performance of these measures when they are computed from predicted score distributions or when they are directly learned from ground-truth data. We find that the latter strategy provides in general better results, though there is still a large space for improvement in aesthetic subjectivity prediction.
Chen Kang, Giuseppe Valenzise, Frédéric Dufaux
MMSP3
2019 A Convex Optimization Framework for Video Quality and Resolution Enhancement From Multiple Descriptions
abstract
Transmission and compression technologies advancement over the past decade led to a shift of multimedia content towards cloud systems. Multiple copies of the same video are available through numerous distribution systems. Different compression levels, algorithms and resolutions are used to match the requirements of particular applications. As 4k display technologies are rapidly adopted, resolution enhancement algorithms are of vital importance. Current solutions do not take into account the particularities of different video encoders, while video reconstruction methods from compressed sources do not provide resolution enhancement. In this paper, we propose a multi source compressed video enhancement framework, where each description can have a different compression level and resolution. Using a variational formulation based on a modern proximal dual splitting algorithm, we efficiently combine multiple descriptions of the same video. Two applications are proposed: combining two compressed low resolution (LR) descriptions of a video sequence into a high resolution (HR) description and enhancing a compressed HR video using a LR compressed description. Tests are performed over multiple video sequences encoded with high efficiency video coding, at different compression levels and resolutions obtained through multiple down-sampling methods.
Andrei I. Purica, Benoit Boyadjis, Béatrice Pesquet-Popescu, Frédéric Dufaux, Cyril Bergeron
IEEE Trans. Image Process.4
2019 Efficient Multi-Strategy Intra Prediction for Quality Scalable High Efficiency Video Coding
abstract
As an extension of High Efficiency Video Coding (HEVC), the Scalable High Efficiency Video Coding (SHVC) introduces multiple layers with inter-layer predictions, which greatly increases the complexity on top of the already complicated HEVC encoder. In Intra prediction for Quality SHVC, Coding Tree Unit (CTU) allows recursive splitting into four depth levels, which considers 35 Intra prediction modes and interlayer reference (ILR) mode to determine the best possible mode at each depth level. This achieves the highest coding efficiency but incurs a substantially high computational complexity. In this paper, we propose a novel Intra prediction scheme to effectively speed up the enhancement layer Intra-coding in Quality SHVC. The new features of the proposed framework include: First, spatial correlation and its correlation degree are combined to predict most probable depth level candidates. Second, for a given depth candidate, based on the probabilities of the ILR mode, we check the ILR mode by examining the residual distribution based on skewness and kurtosis to determine whether the residuals follow a Gaussian distribution. In that case, the Intra prediction comparisons, which require a high complexity, are skipped. Third, during Intra prediction selection from 35 Intra prediction modes, spatial and inter-layer correlations are combined with the local monotonicity of the Hadamard costs associated with the modes in a small neighborhood, to examine only a portion of Intra prediction modes. Finally, a hypothesis testing on the currently selected depth level is performed to examine whether the residuals present significant differences within their block to early terminate depth selection. The proposed multi-step multistrategy scheme aims to minimize the number of depth selections while greatly reducing the mode decision complexity for a depth candidate in a hierarchical fashion. Our experimental results demonstrate that the proposed scheme can achieve a speedup gain of more than 75% in average on the test video sequences, while maintaining almost the same coding efficiency. .
Ce Zhu, Yu Sun 0003, Frédéric Dufaux, Yuanyuan Huang 0007
IEEE Trans. Image Process.4
2019 Learning-Based Tone Mapping Operator for Efficient Image Matching
abstract
In this paper, we propose a new framework to optimally tone map the high dynamic range (HDR) content for image matching under drastic illumination variations. Since tone mapping operators (TMO) have traditionally been used for displaying HDR scenes, their design is suboptimal when used for computer vision tasks, such as image matching. We address this suboptimality by proposing a two-step framework, consisting of: first, a luminance-invariant guidance model based on a support vector regressor (SVR) to optimally adapt the tone mapping function for image matching; and second, an energy maximization model to generate appropriate training samples for learning the SVR. At each step, we collectively address both stages of keypoint detection and descriptor extraction in the feature matching framework. By locally altering the intrinsic characteristics of the tone mapping function, the learned guidance model facilitates the extraction of local invariant features in the presence of illumination variations. We demonstrate that the proposed TMO significantly outperforms perceptually driven state-of-the-art TMOs on a dataset of HDR scenes characterized by challenging lighting variations, such as day/night transitions.
Aakanksha Rana, Giuseppe Valenzise, Frédéric Dufaux
IEEE Trans. Multim.3
2018 Video enhancement with convex optimization methods
abstract
Video enhancement methods enable to optimize the viewing of video content at the end-user side. Most approaches do not consider the compressed nature of the available content. In the present work, we build upon a recently proposed video enhancement approach that explicitly models a compression stage. To apply the enhancement framework on compressed representations requires to extract specific syntax elements during their decoding. This additional information embeds the enhanced result in a domain that closely fits the observation. We evaluate the framework performance in a single source resolution enhancement scenario, and show the method efficiency with respect to state-of-the-art approaches.
Benoit Boyadjis, Andrei I. Purica, Béatrice Pesquet-Popescu, Frédéric Dufaux
ICASSP4
2018 Learning with A Generative Adversarial Network From a Positive Unlabeled Dataset for Image Classification
abstract
In this paper, we propose a new approach which addresses the Positive Unlabeled learning challenge for image classification. Its functioning is based on GAN abilities in order to generate fake images samples whose distribution gets closer to negative samples distribution included in the unlabeled dataset available, while being different to the distribution of the unlabeled positive samples. Then we train a CNN classifier with the positive samples and the fake generated samples, as it would be done with a classic Positive Negative dataset. The tests performed on three different image classification datasets show that the system is stable up to an acceptable fraction of positive samples present in the unlabeled dataset. Although very different, this method outperforms the state of the art PU learning on the RGB dataset CIFAR-10.
Florent Chiaroni, Mohamed-Cherif Rahal, Nicolas Hueber, Frédéric Dufaux
ICIP4
2018 Learning Local Distortion Visibility from Image Quality
abstract
Accurate prediction of local distortion visibility thresholds is critical in many image and video processing applications. Existing methods require an accurate modeling of the human visual system, and are derived through pshycophysical experiments with simple, artificial stimuli. These approaches, however, are difficult to generalize to natural images with complex types of distortion. In this paper, we explore a different perspective, and we investigate whether it is possible to learn local distortion visibility from image quality scores. We propose a convolutional neural network based optimization framework to infer local detection thresholds in a distorted image. Our model is trained on multiple quality datasets, and the results are correlated with empirical visibility thresholds collected on complex stimuli in a recent study. Our results are comparable to state-of-the-art mathematical models that were trained on phsycovisual data directly. This suggests that it is possible to predict psychophysical phenomena from visibility information embedded in image quality scores.
Navaneeth K. Kottayil, Irene Cheng 0001, Giuseppe Valenzise, Frédéric Dufaux
ICIP4
2018 TRISK: A local features extraction framework for texture-plus-depth content matching
Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux
Image Vis. Comput.3
2018 Spatio-temporal constrained tone mapping operator for HDR video compression
Cagri Ozcinar, Paul Lauga, Giuseppe Valenzise, Frédéric Dufaux
J. Vis. Commun. Image Represent.4
2018 Fine-grained detection of inverse tone mapping in HDR images
Wei Fan 0004, Giuseppe Valenzise, Francesco Banterle, Frédéric Dufaux
Signal Process.4
2018 Video analytical coding: When video coding meets video analysis
Ce Zhu, Min Mao, Fangliang Song, Frédéric Dufaux, Xiang Zhang 0006
Signal Process. Image Commun.5
2018 Short-Distance Intra Prediction of Screen Content in Versatile Video Coding (VVC)
abstract
A novel intra prediction algorithm is proposed to improve the coding performance of screen content for the emerging Versatile Video Coding (VVC) standard. The algorithm, called in-loop residual coding with scalar quantization, employs in-block pixels as reference rather than the regular out-block ones. To this end, an additional in-loop residual signal is used to partially reconstruct the block at the pixel level, during the prediction. The proposed algorithm is essentially designed to target high detail textures, where deep block partitioning structure is required. Therefore, it is implemented to operate on 4× 4 blocks only, where further block split is not allowed and the standard algorithm is still unable to properly predict the texture. Experiments in the Joint Exploration Model (JEM) reference software show that the proposed algorithm brings a Bjontegaard Delta (BD)-rate gain of 13% on synthetic content, with a negligible computational complexity overhead at both encoder and decoder sides.
Mohsen Abdoli, Félix Henry, Patrice Brault, Pierre Duhamel, Frédéric Dufaux
IEEE Signal Process. Lett.5
2018 Blind Quality Estimation by Disentangling Perceptual and Noisy Features in High Dynamic Range Images
abstract
High dynamic range (HDR) image visual quality assessment in the absence of a reference image is challenging. This research topic has not been adequately studied largely due to the high cost of HDR display devices. Nevertheless, HDR imaging technology has attracted increasing attention, because it provides more realistic content, consistent to what the human visual system perceives. We propose a new no-reference image quality assessment (NR-IQA) model for HDR data based on convolutional neural networks. The proposed model is able to detect visual artifacts, taking into consideration perceptual masking effects, in a distorted HDR image without any reference. The error and perceptual masking values are measured separately, yet sequentially, and then processed by a mixing function to predict the perceived quality of the distorted image. Instead of using simple stimuli and psychovisual experiments, perceptual masking effects are computed from a set of annotated HDR images during our training process. Experimental results demonstrate that our proposed NR-IQA model can predict HDR image quality as accurately as state-of-the-art full-reference IQA methods.
Navaneeth K. Kottayil, Giuseppe Valenzise, Frédéric Dufaux, Irene Cheng 0001
IEEE Trans. Image Process.3
2017 Good features to track for RGBD images
abstract
RGBD (texture-plus-depth) image representation enriches traditional 2D content with additional geometrical information, having the potential to improve the performance of many computer vision tasks. In image matching, this has been partially studied by considering how depth maps can help render feature descriptors more distinctive. However, little has been done to design keypoint detection approaches able to leverage the availability of depth information. In this paper, we propose a novel and robust approach for detecting corners from RGBD images. Our method modifies a classical corner detection strategy, based on local second-order moment matrices, by computing derivatives in a coordinate system which reflects the local properties of object surfaces. Our results demonstrate a higher stability to out-of-plane rotations of the proposed RGBD corner detector both in terms of feature repeatability and in a visual odometry application.
Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux
ICASSP3
2017 A railroad detection algorithm for infrastructure surveillance using enduring airborne systems
abstract
Infrastructure surveillance is an important requirement for many companies. With the advancement of technology, drones can now provide an efficient tool for such applications. A possible future scenario is the automated surveillance of railroads. Whereas numerous algorithms that provide railroad detection exist, they have mainly focused either on satellite images or for small, low altitude drones which are unsuitable for our particular scenario. In this paper we propose a railroad detection algorithm tailored for large, high altitude enduring drones. More specifically, we use Hough Transform to detect lines and perform a line clustering in the Rho and Theta space. A score model is also proposed in order to identify the railroad. We test our method on several sequences supplied by Airbus Defense & Space and show our algorithm to provide a detection rate of 93.23% in average.
Andrei I. Purica, Béatrice Pesquet-Popescu, Frédéric Dufaux
ICASSP3
2017 Learning-based tone mapping operator for image matching
abstract
In this paper, we propose a new framework to optimally tone-map a high dynamic range (HDR) content for image matching under drastic illumination variations. This task is of fundamental importance for many computer vision applications. To design such a framework, we build a luminance invariant guidance model using a Support Vector Regressor (SVR) and learn it to facilitate the extraction of invariant descriptors from scenes subject to wide variety of appearance changes such as day/night transition. To this end, we initially generate appropriate training samples using a simple similarity-maximization mechanism. We then employ the learned model to predict optimal modulation maps that help to locally alter the intrinsic characteristics (such as shape, size) of the tone mapping function. We evaluate the proposed model performance in terms of matching score and mean average precision rate using state-of-the-art descriptor extraction schemes. We demonstrate that our tone mapping framework significantly outperforms the existing perceptually-driven state-of-the-art TMOs on the benchmark datasets.
Aakanksha Rana, Giuseppe Valenzise, Frédéric Dufaux
ICIP3
2017 Learning-based adaptive tone mapping for keypoint detection
abstract
The goal of tone mapping operators (TMOs) has traditionally been to display high dynamic range (HDR) pictures in a perceptually favorable way. However, when tone-mapped images are to be used for computer vision tasks such as keypoint detection, these design approaches are suboptimal. In this paper, we propose a new learning-based adaptive tone mapping framework which aims at enhancing keypoint stability under drastic illumination variations. To this end, we design a pixel-wise adaptive TMO which is modulated based on a model derived by Support Vector Regression (SVR) using local higher order characteristics. To circumvent the difficulty to train SVR in this context, we further propose a simple detection-similarity-maximization model to generate appropriate training samples using multiple images undergoing illumination transformations. We evaluate the performance of our proposed framework in terms of keypoint repeatability for state-of-the-art keypoint detectors. Experimental results show that our proposed learning-based adaptive TMO yields higher keypoint stability when compared to existing perceptually-driven state-of-the-art TMOs.
Aakanksha Rana, Giuseppe Valenzise, Frédéric Dufaux
ICME3
2017 Intra prediction using in-loop residual coding for the post-HEVC standard
abstract
A few years after standardization of the High Efficiency Video Coding (HEVC), now the Joint Video Exploration Team (JVET) group is exploring post-HEVC video compression technologies. In the intra prediction domain, this effort has resulted in an algorithm with 67 internal modes, new filters and tools which significantly improve HEVC. However, the improved algorithm still suffers from the long distance prediction inaccuracy problem. In this paper, we propose an In-Loop Residual coding Intra Prediction (ILR-IP) algorithm which utilizes inner-block reconstructed pixels as references to reduce the distance from predicted pixels. This is done by using the ILR signal for partially reconstructing each pixel, right after its prediction and before its block-level out-loop residual calculation. The ILR signal is decided in the rate-distortion sense, by a brute-force search on a QP-dependent finite codebook that is known to the decoder. Experiments show that the proposed ILR-IP algorithm improves the existing method in the Joint Exploration Model (JEM) up to 0.45% in terms of bit rate saving, without complexity overhead at the decoder side.
Mohsen Abdoli, Félix Henry, Patrice Brault, Pierre Duhamel, Frédéric Dufaux
MMSP5
2017 Statistical analysis and directional coding of layer-based HDR image coding residue
abstract
Existing methods for layer-based backward compatible high dynamic range (HDR) image and video coding mostly focus on the rate-distortion optimization of base layer while neglecting the encoding of the residue signal in the enhancement layer. Although some recent studies handle residue coding by designing function based fixed global mapping curves for 8-bit conversion and exploiting standard codecs on the resulting 8-bit images, they do not take the local characteristics of residue blocks into account. Inspired by the local anisotropic characteristics of the residue signal and directional methods for motion compensated low dynamic range (LDR) video coding, in this paper we first investigate whether HDR image coding residue exhibits also local anisotropic characteristics. Specifically, we verify directional structures in residue blocks by means of auto-covariance analysis for different bitrates, spatial activities and dynamic ranges as the main variables in HDR image coding. Then, we compare the rate distortion performances of directional coding methods with the baseline residue coding methods in the literature along with different combinations of 8-bit conversion methods. The experiments indicate that content dependent 8-bit conversions and directional coding significantly outperforms the existing function based 8-bit conversions and typical coding for residue coding.
Kutan Feyiz, Fatih Kamisli, Emin Zerman, Giuseppe Valenzise, Alper Koz, Frédéric Dufaux
MMSP6
2017 Quality of experience in UHD-1 phase 2 television: The contribution of UHD+HFR technology
abstract
A key factor to determine the quality of experience (QoE) of a video is its capability to convey the large spectrum of perceptual phenomena that our eyes can sense in real life. In order to meet this demand, the recent DVB UHD-1 Phase 2 specification employs new video features, such as higher spatial resolutions (4K/8K) and High Frame Rate (HFR). The first enables larger field of view and level of details, while the second offers sharper images of moving objects going well beyond the current frame rates. While the contribution of each of these technologies to QoE has been investigated individually, in this paper we are interested to study their interaction, and in quantifying the benefits to users from their combination. To this end, we conduct a subjective test on compressed UHD+HFR content on a recent display capable of reproducing 100 pictures per second at 2160p resolution, with the goal to assess the increase in QoE of UHD and HFR with respect to conventional video, both individually and in combination. The results indicate that for content with fast motion, at higher bitrates the combination of UHD and HFR significantly improves the QoE compared to that obtained when these features are used individually.
Vedad Hulusic, Giuseppe Valenzise, Jean-Charles Gicquel, Jérôme Fournier, Frédéric Dufaux
MMSP5
2017 Analytical distortion aware video coding for computer based video analysis
abstract
With the development of artificial intelligence, more and more multimedia applications for various tasks have emerged in our daily life. Meanwhile, as one of the main information sources of the applications, a huge amount of video data has been being generated by portable or mounted cameras in daily basis for varying purposes including surveillance, in which case we may need computers to "watch" videos to save labor cost. However, most video coding standards are designed for the highest human perceptual quality given a bit rate by minimizing a fidelity cost function (e.g., mean squared error, MSE), assuming the content will be consumed by human beings. In view of the above considerations, this paper proposes a new rate-analytical-distortion optimization method (RADO) for video analysis. Specifically, we consider moving object detection as the analysis task. Accordingly, we develop a novel rate analytical distortion (RAD) model for video coding, where the analytical distortion is related to the object detection performance expressed in terms of F-measure. As shown in the experimental results, the performance of the video analysis task can be significantly improved (up to 40% reduction of analytical distortion) with a slight bit rate increase.
Ce Zhu, Min Mao, Fangliang Song, Frédéric Dufaux, Xiang Zhang 0006
MMSP5
2017 A study of norms in convex optimization super-resolution from compressed sources
abstract
Advancements over the last decade in video acquisition and display technologies lead to a continuous increase of video content resolution. These aspects combined with the shift towards cloud multimedia services and the underway adoption of High Efficiency Video Coding standards (HEVC) create a lot of interest for Super-Resolution (SR) and video enhancing techniques. Recent works showed that proximal based convex optimization approaches provide a promising direction in video restoration. An important aspect in the definition of a SR model is the metric used in defining the objective function. Most techniques are based on the classical I2norm. In this paper we further investigate the use of other norms and their behavior w.r.t. multiple quality evaluation metrics. We show that significant gains of up to 0.5 dB can be obtained when using different norms.
Andrei I. Purica, Benoit Boyadjis, Béatrice Pesquet-Popescu, Frédéric Dufaux
MMSP4
2017 Effect of color space on high dynamic range video compression performance
abstract
High dynamic range (HDR) technology allows for capturing and delivering a greater range of luminance levels compared to traditional video using standard dynamic range (SDR). At the same time, it has brought multiple challenges in content distribution, one of them being video compression. While there has been a significant amount of work conducted on this topic, there are some aspects that could still benefit this area. One such aspect is the choice of color space used for coding. In this paper, we evaluate through a subjective study how the performance of HDR video compression is affected by three color spaces: the commonly used Y'CbCr, and the recently introduced ITP (ICtCp) and Ypu'v'. Five video sequences are compressed at four bit rates, selected in a preliminary study, and their quality is assessed using pairwise comparisons. The results of pairwise comparisons are further analyzed and scaled to obtain quality scores. We found no evidence of ITP improving compression performance over Y'CbCr. We also found that Ypu'v' results in a moderately lower performance for some sequences.
Emin Zerman, Vedad Hulusic, Giuseppe Valenzise, Rafal Mantiuk, Frédéric Dufaux
QoMEX5
2017 AVC to HEVC transcoder based on quadtree limitation
Elie Gabriel Mora, Marco Cagnazzo, Frédéric Dufaux
Multim. Tools Appl.3
2017 A model of perceived dynamic range for HDR images
Vedad Hulusic, Kurt Debattista, Giuseppe Valenzise, Frédéric Dufaux
Signal Process. Image Commun.4
2017 Extended Selective Encryption of H.264/AVC (CABAC)- and HEVC-Encoded Video Streams
abstract
This paper proposes an extended selective encryption (SE) method for both H.264/advanced video coding (AVC) (CABAC) and High Efficiency Video Coding (HEVC) streams, addressing the main security issue that SE is facing: content protection, related to the amount of information leakage through a protected video. Our contribution is the improvement in the visual distortion induced by SE approaches. Previous works on both H.264/AVC (CABAC) and HEVC limit encryption to bins treated by one specific mode of CABAC-its bypass mode-which has the advantage of preserving the overall bitrate, we propose here to also rely on the encryption of the more widely used mode of CABAC-its regular mode. This allows encryption of a major codeword for video reconstruction, the prediction modes for intra blocks/units. Disturbing their statistics may cause bitrate overhead, which is the tradeoff for improving the content security level of the SE approach. A comprehensive study of this compromise between the improvement in the scrambling efficiency and the undesirable aftereffects is presented in this paper, and a specific security analysis of the proposed CABAC regular mode encryption is conducted.
Benoit Boyadjis, Cyril Bergeron, Béatrice Pesquet-Popescu, Frédéric Dufaux
IEEE Trans. Circuits Syst. Video Technol.4
2016 An image smoothing operator for fast and accurate scale space approximation
abstract
Gussian image smoothing is a fundamental operation in the extraction of scale-invariant feature points. Its computation, however, can be too expensive in some resource-constrained scenarios. Alternative solutions such as the box filter can be computed more efficiently, at the cost of a loss in feature repeatibility under some conditions. In this paper we propose a fast and accurate image smoothing operator based on integral images. It has the same order of computational complexity as the box filter, but provides much more accurate visual results and improved keypoint repeatability, which is confirmed in a feature detection scenario using SIFT features.
Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux
ICASSP3
2016 View synthesis based on temporal prediction via warped motion vector fields
abstract
The demand for 3D content has increased over the last years as 3D displays are now widespread. View synthesis methods, such as depth-image-based-rendering, provide an efficient tool in 3D content creation or transmission, and are integrated in coding solutions for multiview video content such as 3D-HEVC. In this paper, we propose a view synthesis method that takes advantage of temporal and inter-view correlations in multiview video sequences. We use warped motion vector fields computed in reference views to obtain temporal predictions of a frame in a synthesized view and blend them with depth-image-based-rendering synthesis. Our method is shown to bring gains of 0.42dB in average when tested on several multiview sequences.
Andrei I. Purica, Marco Cagnazzo, Béatrice Pesquet-Popescu, Frédéric Dufaux, Bogdan Ionescu
ICASSP4
2016 Super-resolution of HEVC videos via convex optimization
abstract
Super Resolution (SR) addresses the problem of image and video upscaling. Most of the best performing SR methods do not take into account any compression prior into the degradation model. Consequently, compression artifacts can be undesirably amplified during SR. In the present work, we propose a novel HEVC-dedicated approach for embedding SR results into a domain that closely fits the compressed observation. Our main contribution is the inclusion of HEVC syntax (block size, quantization parameters etc.) into the degradation model. A recent convex optimization approach is used to solve the associated minimization problem. Over a wide range of resolutions and bitrates, we show that our method improves the results obtained with state of the art SR.
Benoit Boyadjis, Béatrice Pesquet-Popescu, Frédéric Dufaux, Cyril Bergeron
ICIP3
2016 Forensic detection of inverse tone mapping in HDR images
abstract
High dynamic range (HDR) imaging is attracting an increasing deal of attention in the multimedia community, yet its forensic problems have been little studied so far. This paper proposes an HDR image forensic method, which aims at differentiating HDR images created from multiple low dynamic range (LDR) images from those created from a single LDR image by inverse tone mapping. For each kind of HDR image, a Gaussian mixture model is learned. Thereafter, an HDR image forensic feature is constructed based on calculating the Fisher scores. With comparison to a steganalytic feature and a texture/facial analysis feature, experimental results demonstrate the efficiency of the proposed method in HDR image forensic classification on whole images as well as small blocks, for three inverse tone mapping methods.
Wei Fan 0004, Giuseppe Valenzise, Francesco Banterle, Frédéric Dufaux
ICIP4
2016 Background simplification for ROI-oriented low bitrate video coding
abstract
Low-bitrate video compression is a challenging task, particularly with the increasing complexity of video sequences. Re-shaping video data before its compression with modern hybrid encoders has provided interesting results in the low and ultra-low bit rate domains. In this work, we propose a novel saliency guided preprocessing approach, which combines adaptive re-sampling and background texture removal, to achieve efficient ROI-oriented compression. Evaluated with HEVC, we show that our solution improves the ROI encoding over a wide range of resolutions and bit rates whilst maintaining a high background intelligibility level.
Benoit Boyadjis, Cyril Bergeron, Béatrice Pesquet-Popescu, Frédéric Dufaux
MMSP4
2016 Optimizing tone mapping operators for keypoint detection under illumination changes
abstract
Tone mapping operators (TMO) have recently raised interest for their capability to handle illumination changes. However, these TMOs are optimized with respect to perception rather than image analysis tasks like key point detection. Moreover, no work has been done to analyze the factors affecting the optimization of TMOs for such tasks. In this paper, we investigate the influence of two factors-Correlation Coefficient (CC) and Repeatability Rate (RR) of the tone mapped images for the optimization of classical Retinex based models to enhance key point detection under illumination changes. CC-based optimized models aim at increasing the similarity of the tone mapped images. Conversely, RR-based optimized models quantify the optimal detection performance gains. By considering two simple Retinex based models, i.e., Gaussian and bilateral filtering, we show that estimating as precisely as possible the illumination, CC-based optimized models do not necessarily bring to optimal key point detection performance. We conclude that, instead, other criteria specific to RR-based optimized models should be taken into account. Moreover, large gains in performance with respect to existing popular TMOs motivate further research towards optimal tone mapping technique for computer vision applications.
Aakanksha Rana, Giuseppe Valenzise, Frédéric Dufaux
MMSP3
2016 Perceived dynamic range of HDR images
abstract
Although high dynamic range (HDR) imaging has gained great popularity and acceptance in both the scientific and commercial domains, the relationship between perceptually accurate, content-independent dynamic range and objective measures has not been fully explored. In this paper, a new methodology for perceived dynamic range evaluation of complex stimuli in HDR conditions is proposed. A subjective study with 20 participants was conducted and correlations between mean opinion scores (MOS) and three image features were analyzed. Strong Spearman correlations between MOS and objective DR measure and between MOS and image key were found. An exploratory analysis reveals that additional image characteristics should be considered when modeling perceptually-based dynamic range metrics. Finally, one of the outcomes of the study is the perceptually annotated HDR image dataset with MOS values, that can be used for HDR imaging algorithms and metric validation, content selection and analysis of aesthetic image attributes.
Vedad Hulusic, Giuseppe Valenzise, Edoardo Provenzi, Kurt Debattista, Frédéric Dufaux
QoMEX5
2016 Using region-of-interest for quality evaluation of DIBR-based view synthesis methods
abstract
As 3D media became more and more popular over the last years, new technologies are needed in the transmission, compression and creation of 3D content. One of the most commonly used techniques for aiding with the compression and creation of 3D content is known as view synthesis. The most effective class of view synthesis algorithms are using Depth-Image-Based-Rendering techniques, which use explicit scene geometry to render new views. However, these methods may produce geometrical distortions and localized artifacts which are difficult to evaluate as they are inherently different from encoding errors and they are perceived differently by human subjects. In this paper, we propose a region-of-interest evaluation technique for view synthesis based on DIBR methods. Based on the assumption that certain areas determined by the geometrical properties of the scene are prone to distortions, we select a ROI by analyzing the multiple DIBR methods together with the ground truth. The approach is tested using a subjective evaluation view synthesis database and show that our method improves the SSIM correlation with subjective scores We also test another similar method and traditional metrics.
Andrei I. Purica, Giuseppe Valenzise, Béatrice Pesquet-Popescu, Frédéric Dufaux
QoMEX4
2016 An evaluation of HDR image matching under extreme illumination changes
abstract
High dynamic range (HDR) imaging has potential to facilitate computer vision tasks such as image matching where lighting transformations hinder the matching performance. However, little has been done to quantify the gains with different possible HDR representations for vision algorithms like feature extraction. In this paper, we evaluate the performance of the full feature extraction pipeline, including detection and description, on ten different image representations: low dynamic range (LDR), seven different tone mapped (TM) HDR and two HDR imaging (linear and log encoded) representations. We measure the impact of using these different representations for feature matching using mean average precision (mAP) scores on four illumination change datasets. We perform feature extraction using four popular schemes in the literature: SIFT, SURF, BRISK, FREAK. With respect to previous studies, our observations confirm the advantages of HDR over conventional LDR imagery, and the fact that HDR linear values are not appropriate for vision tasks. However, HDR representations that work best for keypoint detection are not necessarily optimal when the full feature extraction is taken into account.
Aakanksha Rana, Giuseppe Valenzise, Frédéric Dufaux
VCIP3
2016 Lagrangian Multiplier Adaptation for Rate-Distortion Optimization With Inter-Frame Dependency
abstract
Rate-distortion optimization (RDO) is widely used in video coding, which plays a critical role in enhancing the coding efficiency substantially. Currently, the RDO process is performed in a way that coding efficiency of each coding unit (CU) is maximized independently without considering the dependency among CUs. As we know, in the current hybrid video coding structure, spatial/temporal prediction techniques are extensively used, which introduce strong dependency among CUs. In this paper, we investigate RDO with inter-frame dependency, where the impact of coding performance of the current CU on that of the following frames is considered. Accordingly, an RDO scheme taking the inter-frame dependency into account is proposed by adapting the Lagrangian multiplier. The experimental results show that the proposed scheme can achieve about 3.22% and 3.19% BD-rate saving in average over the state-of-the-art High Efficiency Video Coding (HEVC) reference software HM15.0 in the low-delay $P$ (LDP) and low-delay $B$ (LDB) coding structures, respectively, with no extra encoding time. The proposed scheme can obtain a significantly higher coding gain than the multiple quantization parameter (MQP) (±3) optimization technique that would greatly increase the encoding time by a factor of about six. Coupled with MQP optimization, the proposed scheme can further achieve about 5.96% and 5.57% BD-rate savings in average over the HEVC and about 4.03% and 4.07% over the HEVC with MQP optimization, under the specified common test conditions for LDP and LDB coding structures, respectively.
Shuai Li 0005, Ce Zhu, Yanbo Gao, Yimin Zhou 0002, Frédéric Dufaux, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.5
2016 Keypoint Detection in RGBD Images Based on an Anisotropic Scale Space
abstract
The increasing availability of texture+depth (RGBD) content has recently motivated research toward the design of image features able to employ the additional geometrical information provided by depth. Indeed, such features are supposed to provide higher robustness than conventional 2D features in the presence of large changes of camera viewpoint. In this paper, we consider the first stage of RGBD image matching, i.e., keypoint detection. In order to obtain viewpoint-covariant keypoints, we design a filtering process, which approximates a diffusion process along the surfaces of the scene, by means of the information provided by depth. Next, we employ this multiscale representation to find keypoints through a multiscale keypoint detector. The keypoints obtained by the proposed detector provide substantially higher stability to viewpoint changes than alternative 2D and RGBD feature extraction approaches, both in terms of repeatability and image classification accuracy. Furthermore, the proposed detector can be efficiently implemented on a GPU.
Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux
IEEE Trans. Multim.3
2015 Improving distinctiveness of brisk features using depth maps
abstract
Binary local descriptors are widely used in computer vision thanks to their compactness and robustness to many image transformations such as rotations or scale changes. However, more complex transformations, like changes in camera viewpoint, are difficult to deal with using conventional features due to the lack of geometric information about the scene. In this paper, we propose a local binary descriptor which assumes that geometric information is available as a depth map. It employs a local parametrization of the scene surface, obtained through depth information, which is used to build a BRISK-like sampling pattern intrinsic to the scene surface. Although we illustrate the proposed method using the BRISK architecture, the obtained parametrization is rather general and could be embedded into other binary descriptors. Our simulations on a set of synthetically generated scenes show that the proposed descriptor is significantly more stable and distinctive than popular BRISK descriptors under a wide range of viewpoint angle changes.
Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux
ICIP3
2015 A scale space for texture+depth images based on a discrete laplacian operator
abstract
In this paper we design a smoothing filter for texture+depth images based on anisotropic diffusion. Our proposed filter enables to generate a scale space on the texture image guided by depth information, and is linear and numerically stable. We show experimentally that using scene geometry preserves the internal structure of 3D surfaces (e.g., it avoids smoothing across object boundaries). As a consequence, the result of smoothing is more independent to changes in the camera position. To illustrate the practical utility of a scale space with such properties, we integrate our filter into the SIFT keypoint detector, getting a substantial improvement of the repeatability of detected keypoints under significant viewpoint position changes.
Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux
ICME3
2015 Inter-frame dependent rate-distortion optimization using lagrangian multiplier adaption
abstract
It is known that, in the current hybrid video coding structure, spatial and temporal prediction techniques are extensively used which introduce strong dependency among coding units. Such dependency poses a great challenge to perform a global rate-distortion optimization (RDO) when encoding a video sequence. RDO is usually performed in a way that coding efficiency of each coding unit is optimized independently without considering dependeny among coding units, leading to a suboptimal coding result for the whole sequence. In this paper, we investigate the inter-frame dependent RDO, where the impact of coding performance of the current coding unit on that of the following frames is considered. Accordingly, an inter-frame dependent rate-distortion optimization scheme is proposed and implemented on the newest video coding standard High Efficiency Video Coding (HEVC) platform. Experimental results show that the proposed scheme can achieve about 3.19% BD-rate saving in average over the state-of-the-art HEVC codec (HM15.0) in the low-delay B coding structure, with no extra encoding time. It obtains a significantly higher coding gain than the multiple QP (±3) optimization technique which would greatly increase the encoding time by a factor of about 6. Coupled with the multiple QP optimization, the proposed scheme can further achieve a higher BD-rate saving of 5.57% and 4.07% in average than the HEVC codec and the multiple QP optimization enabled HEVC codec, respectively.
Shuai Li 0005, Ce Zhu, Yanbo Gao, Yimin Zhou 0001, Frédéric Dufaux, Ming-Ting Sun
ICME5
2015 Foveated High Efficiency Video Coding for Low Bit Rate Transmission
abstract
This work describes the design and subjective performance of Foveated High Efficiency Video Coding (FHEVC). Even though foveation has been widely used for various forms of compression since the early 1990s, we believe its use to improve HEVC is new. We consider the application of, possibly moving, foveated compression in this work and evaluate scenarios where it can be used to improve perceptual quality of videos under constrained transmission resources, e.g., bandwidth. A new method to reduce artifacts during remapping is also proposed. The preliminary implementation considers a single fovea only. Experiments summarizing user evaluations are presented to validate our implementation.
Irene Cheng 0001, Masha Mohammadkhani, Anup Basu, Frédéric Dufaux
ISM4
2015 Evaluation of Feature Detection in HDR Based Imaging Under Changes in Illumination Conditions
abstract
High dynamic range (HDR) imaging enables to capture details in both dark and very bright regions of a scene, and is therefore supposed to provide higher robustness to illumination changes than conventional low dynamic range (LDR) imaging in tasks such as visual features extraction. However, it is not clear how much this gain is, and which are the best modalities of using HDR to obtain it. In this paper we evaluate the first block of the visual feature extraction pipeline, i.e., keypoint detection, using both LDR and different HDR-based modalities, when significant illumination changes are present in the scene. To this end, we captured a dataset with two scenes and a wide range of illumination conditions. On these images, we measure how the repeatability of either corner or blob interest points is affected with different LDR/HDR approaches. Our observations confirm the potential of HDR over conventional LDR acquisition. Moreover, extracting features directly from HDR pixel values is more effective than first tonemapping and then extracting features, provided that HDR luminance information is previously encoded to perceptually linear values.
Aakanksha Rana, Giuseppe Valenzise, Frédéric Dufaux
ISM3
2015 Subjectie evaluation of Super Multi-View compressed contents on high-end light-field 3D displays
Antoine Dricot, Joël Jung, Marco Cagnazzo, Béatrice Pesquet-Popescu, Frédéric Dufaux, Péter Tamás Kovács, Vamsi Kiran Adhikarla
Signal Process. Image Commun.5
2015 Fusion of Global and Local Motion Estimation Using Foreground Objects for Distributed Video Coding
abstract
The side information (SI) in Distributed Video Coding (DVC) is estimated using the available decoded frames and exploited for the decoding and reconstruction of other frames. The quality of the SI has a strong impact on the performance of DVC. Here, we propose a new approach that combines both global and local SI to improve coding performance. Since the background pixels in a frame are assigned to global estimation and the foreground objects to local estimation, one needs to estimate foreground objects in the SI using the backward and forward foreground objects, the background pixels are directly taken from the global SI. Specifically, elastic curves and local motion compensation are used to generate the foreground objects masks in the SI. Experimental results show that, as far as the rate-distortion performance is concerned, the proposed approach can achieve a PSNR improvement of up to 1.39 dB for a group of picture (GOP) size of 2, and up to 4.73 dB for larger GOP sizes, with respect to the reference DISCOVER codec.
Abdalbassir Abou-Elailah, Frédéric Dufaux, Joumana Farah, Marco Cagnazzo, Anuj Srivastava, Béatrice Pesquet-Popescu
IEEE Trans. Circuits Syst. Video Technol.2
2014 Full-reference and reduced-reference quality metrics based on SIFT
abstract
In the last decade, an important research effort has been dedicated to implement objective image quality assessment metrics that reflect effectively human perception. Therefore, the aim of this paper is to propose new objective metrics that fulfill the demands of the image quality assessment field. For this sake, we propose two main full-reference (FR) quality metrics, and then adapt them in such a way to constitute several new reduced-reference (RR) quality metrics, for the case where the complete reference image is not available. We evaluate the influence of five types of distortion such as JPEG, JPEG2000, Gaussian Blur, AWGN, and Contrast change, on the image quality. The proposed metrics are based on the number of Scale-Invariant Feature Transform (SIFT) points, the number of SIFT matches between the unpaired and distorted images, and the Structural Similarity index (SSIM). In order to validate our proposed metrics, we compute the correlation between our metrics' scores and the subjective evaluation results. The results show a high correlation and a better quality range compared to well-known metrics, as well as a good robustness to reduced-reference situations.
Joumana Farah, Marie-Rita Hojeij, Jihad Chrabieh, Frédéric Dufaux
ICASSP4
2014 Full parallax super multi-view video coding
abstract
Super Multi-View (SMV) video is a key enabler for future 3D video services that allows a glasses-free visualization and eliminates many causes of discomfort existing in current available 3D video technologies. SMV video content is composed of tens or hundreds of views, that can be aligned in horizontal only or both horizontal and vertical directions, providing respectively horizontal parallax or full parallax. This paper compares several coding schemes and coding orders, and proposes a coding structure that exploits inter-view correlations in the two directions, providing BD-rate gains up to 29.1% when compared to a basic anchor structure. Additionally, Neighboring Block Disparity Vector (NBDV) and Inter-View Motion Prediction (IVMP) coding tools are further improved to efficiently exploit coding structures in two dimensions, with BD-rate gains up to 4.2% reported over the reference 3D-HEVC encoder.
Antoine Dricot, Joël Jung, Marco Cagnazzo, Béatrice Pesquet-Popescu, Frédéric Dufaux
ICIP5
2014 Local visual features extraction from texture+depth content based on depth image analysis
abstract
With the increasing availability of low-cost - yet precise - depth cameras, “texture+depth” content has become more and more popular in several computer vision and 3D rendering tasks. Indeed, depth images bring enriched geometrical information about the scene which would be hard and often impossible to estimate from conventional texture pictures. In this paper, we investigate how the geometric information provided by depth data can be employed to improve the stability of local visual features under a large spectrum of viewpoint changes. Specifically, we leverage depth information to derive local projective transformations and compute descriptor patches from the texture image. Since the proposed approach may be used with any blob detector, it can be seamlessly integrated into the processing chain of state-of-the-art visual features such as SIFT. Our experiments show that a geometry-aware feature extraction can bring advantages in terms of descriptor distinctiveness with respect to state-of-the-art scale and affine-invariant approaches.
Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux
ICIP3
2014 Methods for improving the tone mapping for backward compatible high dynamic range image and video coding
Alper Koz, Frédéric Dufaux
Signal Process. Image Commun.2
2013 Spatio-temporal saliency based on rare model
abstract
In this paper, a new spatio-temporal saliency model is presented. Based on the idea that both spatial and temporal features are needed to determine the saliency of a video, this model builds upon the fact that locally contrasted and globally rare features are salient. The features used in the model are both spatial (color and orientations) and temporal (motion amplitude and direction) at several scales. To be more robust to moving camera a module computes the global motion and to be more consistent in time, the saliency maps are combined together after a temporal filtering. The model is evaluated on a dataset of 24 videos split into 5 categories (Abnormal, Surveillance, Crowds, Moving camera, and Noisy). This model achieves better performance when compared to several state-of-the-art saliency models.
Marc Décombas, Nicolas Riche, Frédéric Dufaux, Béatrice Pesquet-Popescu, Matei Mancas, Bernard Gosselin, Thierry Dutoit
ICIP3
2013 Optimized tone mapping with LDR image quality constraint for backward-compatible high dynamic range image and video coding
abstract
Backward compatibility to low dynamic range (LDR) displays is an important requirement for high dynamic range (HDR) image and video coding in order to enable a successful transition to HDR technology. In a recent work [1], an optimized solution for tone mapping and inverse tone mapping of HDR images is achieved in terms of mean square error (MSE) of the logarithm of luminance values of HDR image pixels for backward-compatible compression. Although this pioneer optimization approach provides a well settled mathematical framework for tone mapping, one of its important shortcomings is not to take the quality of the resulting LDR images into account during the formulation. In this paper, we include the LDR image quality as a constraint to optimization problem and develop a methodology to compromise the trade-off between HDR image quality and LDR image quality during HDR image and video coding. The developed methodology is verified on HDR images by showing the increase (decrease) in the quality of generated LDR images while losing (gaining) from the rate-distortion performance of HDR image coding.
Alper Koz, Frédéric Dufaux
ICIP2
2013 Comparison of DASH adaptation strategies based on bitrate and quality signalling
abstract
During a Dynamic Adaptive Streaming over HTTP (DASH) video transmission, the client can dynamically adapt the bitrate of the requested stream to changes in network conditions, by switching between different versions of the same content encoded at different bitrates, called representations. Each representation is split in smaller portions, called segments, and representation switching can be done at each segment. In this paper, we propose a comparative analysis of different adaptation strategies for representation switching, which exploit information on the bitrate and the visual quality of the different available versions of the requested content, at representation or at segment level. The results allow to identify the advantages and drawbacks related to each strategy in terms of bandwidth exploitation, maximization of visual quality and risk of buffer underflow.
Francesca De Simone, Frédéric Dufaux
MMSP2
2013 Segmentation-based optimized tone mapping for high dynamic range image and video coding
abstract
A core part of the state-of-the art high dynamic range (HDR) image and video compression methods is the tone mapping operation to convert the visible luminance range into the finite bit depths that can be supported by the current video codecs. These conversions are until now optimized to provide backward compatibility to the existing low dynamic range (LDR) displays. However, a direct application of these methods for the emerging HDR displays can result in a loss of details in the bright and dark regions of the HDR content. In this paper, we overcome this limitation by designing a tone mapping operation which handles the bright and dark regions separately. The proposed method first finds the optimal segmentation of the HDR image into two parts, namely dark and bright regions, and then designs the optimal tone mapping for each region in terms of the mean square error between the logarithm of the luminance values of the original and reconstructed HDR content (HDR-MSE). The results indicate the superiority of the proposed method over the state-of-the art HDR coding methods.
Paul Lauga, Alper Koz, Giuseppe Valenzise, Frédéric Dufaux
PCS4
2013 Fusion of Global and Local Motion Estimation for Distributed Video Coding
abstract
The quality of side information plays a key role in distributed video coding. In this paper, we propose a new approach that consists of combining global and local motion compensation at the decoder side. The parameters of the global motion are estimated at the encoder using scale invariant feature transform features. Those estimated parameters are sent to the decoder in order to generate a globally motion compensated side information. Conversely, a locally motion compensated side information is generated at the decoder based on motion-compensated temporal interpolation of neighboring reference frames. Moreover, an improved fusion of global and local side information during the decoding process is achieved using the partially decoded Wyner-Ziv frame and decoded reference frames. The proposed technique improves significantly the quality of the side information, especially for sequences containing high global motion. Experimental results show that, as far as the rate-distortion performance is concerned, the proposed approach can achieve a PSNR improvement of up to 1.9 dB for a Group of Pictures (GOP) size of 2, and up to 4.65 dB for larger GOP sizes, with respect to the reference DISCOVER codec.
Abdalbassir Abou-Elailah, Frédéric Dufaux, Joumana Farah, Marco Cagnazzo, Béatrice Pesquet-Popescu
IEEE Trans. Circuits Syst. Video Technol.2
2013 Guest Editorial: Special issue on intelligent video surveillance for public security and personal privacy
abstract
This Special Issue offers an overview of ongoing research on intelligent video surveillance (IVS) techniques, and brings together cutting-edge research work on security and privacy problems with respect to technological, behavioral, legal, and cultural aspects. We received 34 submissions and each submission was rigorously reviewed by at least two experts in the related fields based on the criteria of originality, significance, quality, and clarity. Eventually, 12 papers were accepted for the Special Issue, spanning a variety of topics including privacy protection, background modeling, tracking, action/activity analysis, and crowd behavior perception. The papers constituting this issue are then briefly summarized.
Noboru Babaguchi, Andrea Cavallaro, Rama Chellappa, Frédéric Dufaux, Liang Wang 0001
IEEE Trans. Inf. Forensics Secur.4
2012 A new object based quality metric based on SIFT and SSIM
abstract
We propose a full reference visual quality metric to evaluate a semantic coding system which may not preserve exactly the position and/or the shape of objects. The metric is based on Scale-Invariant Feature Transform (SIFT) points. More specifically, Structural SIMilarity (SSIM) on windows around the SIFT points measures the compression artifacts (SSIM_SIFT). Conversely, the standard deviation of the matching distance between the SIFT points measures the geometric distortion (GEOMETRIC_SIFT). We validate our metric with subjective evaluation and reach a Spearman correlation of 0.86 for SSIM_SIFT and 0.74 for GEOMETRIC_SIFT.
Marc Décombas, Frédéric Dufaux, Erwann Renan, Béatrice Pesquet-Popescu, François Capman
ICIP2
2012 Improved seam carving for semantic video coding
abstract
Traditional video codecs like H.264/AVC encode video sequences to minimize the Mean Squared Error (MSE)at a given bitrate. Seam carving is a content-aware resizing method. In this paper, we propose a semantic video compression scheme based on seam carving. Its principle is to suppress non salient parts of the video by seam carving. The reduced sequence is then encoded with H.264/AVC and the seams are represented and encoded with our proposed approach. The main idea is to encode the seams by regrouping them. Compared to our earlier work, the main contributions of this paper are: a new energy map with better temporal robustness, a new way to define groups of seams using k-median clustering, and an improved background synthesis. Experiments show that, compared to a traditional H.264/AVC encoding, we reach a bitrate saving between 10% and 24%%with the same quality of the salient objects.
Marc Décombas, Frédéric Dufaux, Erwann Renan, Béatrice Pesquet-Popescu, François Capman
MMSP2
2012 Optimized tone mapping with perceptually uniform luminance values for backward-compatible high dynamic range video compression
abstract
Backward compatibility for high dynamic range image and video compression forms one of the essential requirements in the transition phase from low dynamic range (LDR) displays to high dynamic range (HDR) displays. In a recent work [1], an optimized solution for tone mapping and inverse tone mapping of HDR images is achieved in terms of mean square error (MSE) of the logarithm of luminance values of HDR image pixels for backward-compatible compression. A disadvantage of this approach was to use non uniform luminance values according to Human perception for minimization, which causes quite non-natural over-illumination in the produced LDR images. In this paper, we propose to use perceptually uniform luminance values as an alternative for the optimization of tone mapping curve. The results indicate that the proposed approach gives better performance (0.5–1 dB gains) in terms of Perceptually Uniform Peak Signal to Noise Ratio (PU-PSNR) and produces more realistic LDR images.
Alper Koz, Frédéric Dufaux
VCIP2
2012 JPSearch: New international standard providing interoperable framework for image search and sharing
Kyoungro Yoon, Youngseop Kim, Je-Ho Park, Jaime Delgado, Akio Yamada, Frédéric Dufaux, Rubén Tous
Signal Process. Image Commun.6
2011 Using distributed source coding and depth image based rendering to improve interactive multiview video access
abstract
Multiple-views video is commonly believed to be the next significant achievement in video communications, since it enables new exciting interactive services such as free viewpoint television and immersive teleconferencing. However the interactivity requirement (i.e. allowing the user to change the viewpoint during video streaming) involves a trade-off between storage and bandwidth costs. Several solutions have been proposed in the literature, using redundant predictive frames, Wyner-Ziv frames, or a combination of them. In this paper, we adopt distributed video coding for interactive multiview video plus depth (MVD), taking advantage of depth image based rendering (DIBR) and depth-aided inpainting to fill the occlusion areas. To the authors' best knowledge, very few works in interactive MVD consider the problem of continuity of the playback during the switching among streams. Therefore we survey the existing solutions, we propose a set of techniques for MVD coding and we compare them. As main results, we observe that DIBR can help in rate reduction (up to 13.36% for the texture video and up to 8.67% for the depth map, wrt the case where DIBR is not used), and we also note that the optimal strategy to combine DIBR and distributed video coding depends on the position of the switching time into the group of pictures. Choosing the best technique on a frame-to-frame basis can further reduce the rate from 1% to 6%.
Giovanni Petrazzuoli, Marco Cagnazzo, Frédéric Dufaux, Béatrice Pesquet-Popescu
ICIP3
2011 Wyner-ziv coding for depth maps in multiview video-plus-depth
abstract
Three dimensional digital video services are gathering a lot of attention in recent years, thanks to the introduction of new and efficient acquisition and rendering devices. In particular, 3D video is often represented by a single view and a so called depth map, which gives information about the distance between the point of view and the objects. This representation can be extended to multiple views, each with its own depth map. Efficient compression of this kind of data is of course a very important topic in sight of a massive deployment of services such as 3D-TV and FTV (free viewpoint TV). In this paper we consider the application of distributed coding techniques to the coding of depth maps, in order to reduce the complexity of single view or multi view encoders and to enhance interactive multiview video streaming. We start from state-of-the-art distributed video coding techniques and we improve them by using high order motion interpolation and by exploiting texture motion information to encode the depth maps. The experiments reported here show that the proposed method achieves a rate reduction up to 11.06% compared to state-of-the-art distributed video coding technique.
Giovanni Petrazzuoli, Marco Cagnazzo, Frédéric Dufaux, Béatrice Pesquet-Popescu
ICIP3
2010 A framework for the validation of privacy protection solutions in video surveillance
abstract
The issue of privacy protection in video surveillance has drawn a lot of interest lately. However, thorough performance analysis and validation is still lacking, especially regarding the fulfillment of privacy-related requirements. In this paper, we put forward a framework to assess the capacity of privacy protection solutions to hide distinguishing facial information and to conceal identity. We then conduct rigorous experiments to evaluate the performance of face recognition algorithms applied to images altered by privacy protection techniques. Results show the ineffectiveness of naïve privacy protection techniques such as pixelization and blur. Conversely, they demonstrate the effectiveness of more sophisticated scrambling techniques to foil face recognition.
Frédéric Dufaux, Touradj Ebrahimi
ICME1
2010 Encoder and decoder side global and local motion estimation for Distributed Video Coding
abstract
In this paper, we propose a new Distributed Video Coding (DVC) architecture where motion estimation is performed both at the encoder and decoder, effectively combining global and local motion models. We show that the proposed approach improves significantly the quality of Side Information (SI), especially for sequences with complex motion patterns. In turn, it leads to rate-distortion gains of up to 1 dB when compared to the state-of-the-art DISCOVER DVC codec.
Frédéric Dufaux, Touradj Ebrahimi
MMSP1
2010 Special issue on Image and Video Quality Assessment
Stefan Winkler 0001, Frédéric Dufaux, Dominique Barba, Vittorio Baroncini
Signal Process. Image Commun.2
2009 Towards Generic Detection of Unusual Events in Video Surveillance
abstract
In this paper, we consider the challenging problem of unusual event detection in video surveillance systems. The proposed approach makes a step toward generic and automatic detection of unusual events in terms of velocity and acceleration. At first, the moving objects in the scene are detected and tracked. A better representation of moving objects trajectories is then achieved by means of appropriate pre-processing techniques. A supervised support vector machine method is then used to train the system with one or more typical sequences, and the resulting model is then used for testing the proposed method with other typical sequences (different scenes and scenarios). Experimental results are shown to be promising. The presented approach is capable of determining similar unusual events as in the training sequences.
Frédéric Dufaux, Thien M. Ha, Touradj Ebrahimi
AVSS2
2009 Error-resilient scalable compression based on distributed video coding
Mourad Ouaret, Frédéric Dufaux, Touradj Ebrahimi
Signal Process. Image Commun.2
2008 H.264/AVC video scrambling for privacy protection
abstract
In this paper, we address the problem of privacy in video surveillance systems. More specifically, we consider the case of H.264/AVC which is the state-of-the-art in video coding. We assume that regions of interest (ROI), containing privacy-sensitive information, have been identified. The content of these regions are then concealed using scrambling. More specifically, we introduce two region-based scrambling techniques. The first one pseudo-randomly flips the sign of transform coefficients during encoding. The second one is performing a pseudo-random permutation of transform coefficients in a block. The flexible macroblock ordering (FMO) mechanism of H.264/AVC is exploited to discriminate between the ROI which are scrambled and the background which remains clear. Experimental results show that both techniques are able to effectively hide private information in ROI, while the scene remains comprehensible. Furthermore, the loss in coding efficiency stays small, whereas the required additional computational complexity is negligible.
Frédéric Dufaux, Touradj Ebrahimi
ICIP1
2008 Improved side information generation with iterative decoding and frame interpolation for Distributed Video Coding
abstract
Distributed Video Coding (DVC) is a new paradigm in video coding, which is receiving a lot of interests nowadays. Side Information (SI) generation is a key function in the DVC decoder, and plays a key-role in determining the performance of the codec. This paper proposes an improved side information generation scheme, which exploits both spatial and temporal correlations in the sequences. Partially decoded Wyner-Ziv (WZ) frames, based on initial SI by Motion Compensation Temporal Interpolation (MCTI), are exploited to improve the performance of the whole SI generation. In addition, an enhanced temporal frame interpolation is proposed, including motion vector refinement and smoothing, optimal compensation mode selection, and a new matching criterion for motion estimation. Simulation results show that the proposed scheme can achieve up to 2.3 dB improvement in Rate Distortion (RD) performance for video with high motion, when compared to state-of-the-art DVC.
Shuiming Ye, Mourad Ouaret, Frédéric Dufaux, Touradj Ebrahimi
ICIP3
2008 Hybrid spatial and temporal error concealment for distributed video coding
abstract
Distributed video coding (DVC) is based on a new paradigm in coding, which has received many interests recently. This paper proposes a hybrid spatial and temporal error concealment scheme to conceal errors in Wyner-Ziv (WZ) frames. We first use a spatial concealment based on edge directed filter. This step is exploited to improve the performance of subsequent temporal concealment. An enhanced temporal concealment based on motion compensated temporal interpolation is proposed, including motion vector refinement and smoothing, optimal compensation mode selection, and a new matching criterion for motion estimation. Simulation results show that the objective qualities as well as the perceptual qualities of the corrupted sequences are significantly improved by the hybrid error concealment, outperforming both spatial and temporal concealments alone.
Shuiming Ye, Mourad Ouaret, Frédéric Dufaux, Touradj Ebrahimi
ICME3
2008 Scrambling for Privacy Protection in Video Surveillance Systems
abstract
In this paper, we address the problem of privacy protection in video surveillance. We introduce two efficient approaches to conceal regions of interest (ROIs) based on transform-domain or codestream-domain scrambling. In the first technique, the sign of selected transform coefficients is pseudorandomly flipped during encoding. In the second method, some bits of the codestream are pseudorandomly inverted. We address more specifically the cases of MPEG-4 as it is today the prevailing standard in video surveillance equipment. Simulations show that both techniques successfully hide private data in ROIs while the scene remains comprehensible. Additionally, the amount of noise introduced by the scrambling process can be adjusted. Finally, the impact on coding efficiency performance is small, and the required computational complexity is negligible.
Frédéric Dufaux, Touradj Ebrahimi
IEEE Trans. Circuits Syst. Video Technol.1
2007 Codec-Independent Scalable Distributed Video Coding
abstract
In this paper, we introduce novel schemes for scalable distributed video coding (DVC), dealing with temporal, spatial and quality scalabilities. More specifically, conventional coding is used to obtain a base layer. DVC is then applied to generate enhancement layers. The side information is generated either temporally by motion compensated interpolation, or spatially by a spatial bi-cubic interpolation. Note that this scalable DVC approach is independent from the codec used to encode or decode the base layer. Simulation results show that most of the proposed schemes outperform non-scalable DVC, in addition to enabling the scalability features.
Mourad Ouaret, Frédéric Dufaux, Touradj Ebrahimi
ICIP (3)2
2006 A Novel Replica Detection System using Binary Classifiers, R-Trees, and PCA
abstract
Replica detection is a prerequisite for the discovery of copyright infringement and detection of illicit content. For this purpose, content-based systems can be an efficient alternative to watermarking. Rather than imperceptibly embedding a signal, content-based systems rely-on image similarity. Certain content-based systems use adaptive classifiers to detect replicas. In such systems, a suspect image is tested against every original, which can become computationally prohibitive as the number of original images grows. In this paper, we propose using R-tree indexing to decrease the necessary number of comparisons and rapidly select the most likely originals. Experimental results show that the proposed system performs very satisfactorily and that up to 99.3% of the originals can be discarded before applying the binary classifiers.
Yannick Maret, Spiros Nikolopoulos, Frédéric Dufaux, Touradj Ebrahimi, Nikos Nikolaidis 0001
ICIP3
2006 The emerging JPEG-2000 security (JPSEC) standard
abstract
The emergence of digital imaging applications is accelerating the need for security of digital imagery. The emerging international standard ISO/IEC JPEG-2000 security (JPSEC) is designed to provide security for digital imagery, and in particular digital imagery coded with the JPEG-2000 image coding standard. This paper provides an overview of the JPSEC standard, including a description of its basic architecture and examples of its use.
John G. Apostolopoulos, Susie J. Wee, Frédéric Dufaux, Touradj Ebrahimi, Qibin Sun, Zhishou Zhang
ISCAS3
2006 JPWL - an extension of JPEG 2000 for wireless imaging
abstract
In this paper, we present an overview of the JPWL standardization activity. JPWL is an extension of JPEG 2000 for the efficient transmission of JPEG 2000 images over an error-prone wireless network. More specifically, JPWL supports a set of tools for error protection and correction, including forward error correcting codes (FEC), unequal error protection (UEP), data partitioning and interleaving
Frédéric Dufaux, Giuseppe Baruffa, Fabrizio Frescura, Didier Nicholson
ISCAS1
2006 Adaptive image replica detection based on support vector classifiers
Yannick Maret, Frédéric Dufaux, Touradj Ebrahimi
Signal Process. Image Commun.2
2004 Error-resilient video coding performance analysis of motion JPEG2000 and MPEG-4
abstract
The new Motion JPEG 2000 standard is providing with some compelling features. It is based on an intra-frame wavelet coding, which makes it very well suited for wireless applications. Indeed, the state-of-the-art wavelet coding scheme achieves very high coding efficiency. In addition, Motion JPEG 2000 is very resilient to transmission errors as frames are coded independently (intra coding). Furthermore, it requires low complexity and introduces minimal coding delay. Finally, it supports very efficient scalability. In this paper, we analyze the performance of Motion JPEG 2000 in error-prone transmission. We compare it to the well-known MPEG-4 video coding scheme, in terms of coding efficiency, error resilience and complexity. We present experimental results which show that Motion JPEG 2000 outperforms MPEG-4 in the presence of transmission errors.
Frédéric Dufaux, Touradj Ebrahimi
VCIP1
2004 Perceptual blur and ringing metrics: application to JPEG2000
Pina Marziliano, Frédéric Dufaux, Stefan Winkler 0001, Touradj Ebrahimi
Signal Process. Image Commun.2
2003 Video quality evaluation for mobile streaming applications
Stefan Winkler 0001, Frédéric Dufaux
VCIP2
2002 A no-reference perceptual blur metric
abstract
We present a no-reference blur metric for images and video. The blur metric is based on the analysis of the spread of the edges in an image. Its perceptual significance is validated through subjective experiments. The novel metric is near real-time, has low computational complexity and is shown to perform well over a range of image content. Potential applications include optimization of source coding, network resource management and autofocus of an image capturing device.
Pina Marziliano, Frédéric Dufaux, Stefan Winkler 0001, Touradj Ebrahimi
ICIP (3)2
2001 Constrained bit-rate control for very low bit-rate streaming-video applications
abstract
We propose a low-complexity frame-layer bit-rate control algorithm for very low bit-rate streaming-video applications with causal one-pass processing. Rate control is achieved by jointly adapting the frame rate and quantization step-size. Constraint coding is introduced, which forces the encoder to operate within a subset of operating points on a 2-D grid. The control parameters are also constrained to change in a gradual fashion, allowing bits to be saved in easy scenes so that more bits can be used for difficult scenes. A simple scene-change detector is used to insert intra frames at scene-change boundaries. The proposed algorithm is implemented in H.263 and is compared with the case of fixed control parameters and a conventional off-line bit-rate control based on adapting the quantization step-size. It is shown that the proposed algorithm codes more frames over a given time interval while achieving average PSNR gains over 1 dB and reducing bit-rate fluctuations. Results are confirmed by subjective tests, which show that our bit-rate control consistently provides improved visual quality.
Eric C. Reed, Frédéric Dufaux
IEEE Trans. Circuits Syst. Video Technol.2
2000 Key Frame Selection to Represent a Video
abstract
This paper describes a technique to automatically extract a single key frame from a video sequence. The technique is designed for a system to search video on the World Wide Web. For each video returned by a query, a thumbnail image that illustrates its content is displayed to summarize the results. The proposed technique is composed of three steps. Shot boundaries detection, shot selection, and key frame extraction within the selected shot. The shot and key frame are selected based on measures of motion and spatial activity and the likeliness to include people. The latter is determined by skin-color detection and face detection. Simulation results on a large set of video from the Internet, including movie trailers, sports, news, and animation, show the efficiency of the method. Furthermore, this is achieved at a very low complexity cost.
Frédéric Dufaux
ICIP1
2000 Combined Spline- and Block-Based Motion Estimation for Video Coding
abstract
We propose a technique to estimate motion for video coding. Our technique combines spline-based registration and block matching motion estimation. It replaces the initial compute-intensive gross block matching search with spline-based registration, which is more efficient in recovering full-image motion fields. This results in a smooth motion field representative of the true motion in the scene, which can be more efficiently encoded and be guaranteed of high visual quality as well. Furthermore, the implementation is computationally cost-effective. Experimental results on well-known test image sequences show that the method results in an increase in coding efficiency along with a reduction in computational cost.
Frédéric Dufaux, Sing Bing Kang
ICPR1
2000 Efficient, robust, and fast global motion estimation for video coding
abstract
In this paper, we propose an efficient, robust, and fast method for the estimation of global motion from image sequences. The method is generic in that it can accommodate various global motion models, from a simple translation to an eight-parameter perspective model. The algorithm is hierarchical and consists of three stages. In the first stage, a low-pass image pyramid is built. Then, an initial translation is estimated with full-pixel precision at the top of the pyramid using a modified n-step search matching. In the third stage, a gradient descent is executed at each level of the pyramid starting from the initial translation at the coarsest level. Due to the coarse initial estimation and the hierarchical implementation, the method is very fast. To increase robustness to outliers, we replace the usual formulation based on a quadratic error criterion with a truncated quadratic function. We have applied the algorithm to various test sequences within an MPEG-4 coding system. From the experimental results we conclude that global motion estimation provides significant performance gains for video material with camera zoom and/or pan. The gains result from a reduced prediction error and a more compact representation of motion. We also conclude that the robust error criterion can introduce additional performance gains without increasing computational complexity.
Frédéric Dufaux, Janusz Konrad
IEEE Trans. Image Process.1
1997 Pre and Post-Filtering for Low Bit-Rate Video Coding
abstract
We propose pre and post-filters for low bit rate video coding. The purpose of the former is to make the video sequence easier to encode, whereas the latter aims to remove coding artifacts. The proposed techniques are computationally efficient and lead to scalable architectures which (1) can handle varied types of video content and (2) can be adapted to the available bandwidth and computational resources. Simulation results using H.263 show that significant gains can be achieved by pre and post-filtering.
Nuno Vasconcelos, Frédéric Dufaux
ICIP (1)2
1996 Object tracking based on temporal and spatial information
abstract
This paper addresses the problem of segmenting an image sequence in terms of multiple moving objects and tracking them through time and presents an object tracking algorithm. The objects are characterized through their temporal and spatial features so as to identify them and carry out the tracking procedure. The proposed tracking algorithm helps to detect the objects present in the current frame by supplying previous spatio-temporal information to the spatio-temporal segmentation procedure. In addition, the proposed algorithm tackles the correspondence problem. This is achieved through the use of a multiple hypotheses framework, the latter tests being based on both temporal and spatial characterizations of the objects.
Fabrice Moscheni, Frédéric Dufaux, Murat Kunt
ICASSP2
1996 Background mosaicking for low bit rate video coding
abstract
This paper proposes a new technique to build a background memory based on mosaicking. More precisely, the technique first identifies background and foreground regions based on local motion estimates. Camera motion is then estimated on the background by applying parametric global motion estimation. Finally, after compensating for camera motion, the background content is temporally integrated in long-term memory. The method leads to high coding performances and allows for content-based functionalities.
Frédéric Dufaux, Fabrice Moscheni
ICIP (1)1
1995 A new two-stage global/local motion estimation based on a background/foreground segmentation
abstract
In video coding, the reduction of temporal redundancy is the key to achieving high performance. In the framework of sequence coding, motion estimation and compensation has been shown to be very efficient at removing temporal redundancy. The motion existing in a scene can be mainly seen as arising from local motions superimposed to the camera motion. A new two stage global/local motion estimation approach is presented. The global motion estimation only relies on the background information. It is based on a matching technique and the global motion model is chosen to be affine. Simulation results show significant improvements obtained with the proposed method compared to the usual methods.
Fabrice Moscheni, Frédéric Dufaux, Murat Kunt
ICASSP2
1995 Spatio-temporal segmentation based on motion and static segmentation
abstract
The problem of segmenting an image sequence in terms of regions characterized by a coherent motion is among the most challenging in image sequence analysis. This paper proposes a new technique which sequentially refines the segmentation and the motion estimation by combining static segmentation and motion information. The motion is robustly computed by a global estimation which remove the camera motion, followed by a local estimation using a matching technique and a robust estimator. Simulation results show the efficiency of the proposed technique.
Frédéric Dufaux, Fabrice Moscheni, Andy Lippman
ICIP1
1995 Motion estimation techniques for digital TV: a review and a new contribution
abstract
The key to high performance in image sequence coding lies in an efficient reduction of the temporal redundancies. For this purpose, motion estimation and compensation techniques have been successfully applied. This paper studies motion estimation algorithms in the context of first generation coding techniques commonly used in digital TV. In this framework, estimating the motion in the scene is not an intrinsic goal. Motion estimation should indeed provide good temporal prediction and simultaneously require low overhead information. More specifically the aim is to minimize globally the bandwidth corresponding to both the prediction error information and the motion parameters. This paper first clarifies the notion of motion, reviews classical motion estimation techniques, and outlines new perspectives. Block matching techniques are shown to be the most appropriate in the framework of first generation coding. To overcome the drawbacks characteristic of most block matching techniques, this paper proposes a new locally adaptive multigrid block matching motion estimation technique. This algorithm has been designed taking into account the above aims. It leads to a robust motion field estimation precise prediction along moving edges and a decreased amount of side information in uniform areas. Furthermore, the algorithm controls the accuracy of the motion estimation procedure in order to optimally balance the amount of information corresponding to the prediction error and to the motion parameters. Experimental results show that the technique results in greatly enhanced visual quality and significant saving in terms of bit rate when compared to classical block matching techniques.>
Frédéric Dufaux, Fabrice Moscheni
Proc. IEEE1
1994 A Motion Field Segmentation to Improve Moving Edges Reconstruction in Video Coding
abstract
Block matching motion estimation techniques have shown their efficiency to reduce the temporal redundancy for video coding. However, they tend to introduce annoying block artifacts in the displaced frame difference due to their block-based model. In this paper, we propose a motion field refinement technique in order to overcome these artifacts. The method relies on vector quantization. The patterns to segment the motion field are derived by the LBG algorithm, the training being performed on segmented patterns obtained from natural images. Simulation results show a significant improvement due to the proposed method, in particular moving edges are much better motion compensated and the visual quality is greatly enhanced.>
Iole Moccagatta, Fabrice Moscheni, Markus Schütz, Frédéric Dufaux
ICIP (3)4
1994 Vector Quantization-Based Motion Field Segmentation under the Entropy Criterion
Frédéric Dufaux, Iole Moccagatta, Fabrice Moscheni, Henri Nicolas
J. Vis. Commun. Image Represent.1
1994 Multigrid block matching motion estimation for generic video coding
Frédéric Dufaux
Signal Process.1
1993 Image sequence coding by multigrid motion estimation and segmentation-based coding of prediction errors
abstract
This paper presents an innovative video coding system which processes appropriately motion information by an improved motion estimation algorithm and a segmentation based coding of displaced frame differences (DFD). The proposed multigrid motion estimation algorithm with adaptive mesh leads to more uniform and accurate motion vectors, and a lower overhead information. A morphological segmentation algorithm is proposed to code DFD by sending contours and quantized high energy regions. The method is coding oriented, and has the potential of graceful degradation. Simulation results show outstanding performances of the proposed codec for image sequences in CCIR 601 format.
Frédéric Dufaux
VCIP2
1993 Entropy criterion for optimal bit allocation between motion and prediction error information
abstract
Motion estimation and compensation techniques are widely used in video coding. This paper addresses the problem of the trade-off between the motion and the prediction error information. Under some realistic hypotheses, the transmission cost of these two components can be estimated. Therefore, we obtain a criterion which controls the motion estimation process in order to optimize its performance. As a particular application, this criterion is applied to the split procedure of an adaptive multigrid block matching technique. Simulation results are presented, showing the significant improvements due to the method.
Fabrice Moscheni, Frédéric Dufaux, Henri Nicolas
VCIP2