VLDB 2026 Research / reviewers in the wild / expert
Giuseppe Valenzise
dblp:22/4971
· DBLP profile ↗
93ranked-venue papers
12as first author
34since 2021 · last 2026
0000-0002-5840-5743ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 90 · 11 first-author · 33 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CutClean: Neural Network Pruning for Privacy-Preserving Inference
Leonardo Magliolo, Vito Paolo Pastore, Giuseppe Valenzise, Enzo Tartaglione |
ICPR (11) | 3 |
| 2026 | Compression in 3D Gaussian Splatting: A Survey of Methods, Trends, and Future Directionsabstract3D Gaussian Splatting (3DGS) has recently emerged as a pioneering approach in explicit scene rendering and computer graphics. Unlike traditional neural radiance field (NeRF) methods, which typically rely on implicit, coordinate-based models to map spatial coordinates to pixel values, 3DGS utilizes millions of learnable 3D Gaussians. Its differentiable rendering technique and inherent capability for explicit scene representation and manipulation positions 3DGS as a potential game-changer for the next generation of 3D reconstruction and representation technologies. This enables 3DGS to deliver real-time rendering speeds while offering unparalleled editability levels. However, despite its advantages, 3DGS suffers from substantial memory and storage requirements, posing challenges for deployment on resource-constrained devices. In this survey, we provide a comprehensive overview focusing on the scalability and compression of 3DGS. We begin with a detailed background overview of 3DGS, followed by a structured taxonomy of existing compression methods. Additionally, we analyze and compare current methods from the topological perspective, evaluating their strengths and limitations in terms of fidelity, compression ratios, and computational efficiency. Furthermore, we explore how advancements in efficient NeRF representations can inspire future developments in 3DGS optimization. Finally, we conclude with current research challenges and highlight key directions for future exploration. Muhammad Salman Ali, Chaoning Zhang, Marco Cagnazzo, Giuseppe Valenzise, Enzo Tartaglione, Sung-Ho Bae |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | RAM-VQA: Restoration Assisted Multi-Modality Video Quality AssessmentabstractVideo Quality Assessment (VQA) strives to computationally emulate human perceptual judgments and has garnered significant attention given its widespread applicability. However, existing methodologies face two primary impediments: (1) limited proficiency in evaluating samples at quality extremes (e.g., severely degraded or near-perfect videos), and (2) insufficient sensitivity to nuanced quality variations arising from a misalignment with human perceptual mechanisms. Although vision-language models offer promising semantic understanding, their reliance on visual encoders pre-trained for high-level tasks often compromises their sensitivity to low-level distortions. To surmount these challenges, we propose the Restoration-Assisted Multi-modality VQA (RAM-VQA) framework. Uniquely, our approach leverages video restoration as a proxy to explicitly model distortion-sensitive features. The framework operates through two synergistic stages: a prompt learning stage that constructs a quality-aware textual space using triple-level references (degraded, restored, and pristine) derived from the restoration process, and a dual-branch evaluation stage that integrates semantic cues with technical quality indicators via spatio-temporal differential analysis. Extensive experiments demonstrate that RAM-VQA achieves state-of-the-art performance across diverse benchmarks, exhibiting superior capability in handling extreme-quality content while ensuring robust generalization. Pengfei Chen 0003, Jiebin Yan, Rajiv Soundararajan, Giuseppe Valenzise, Leida Li |
IEEE Trans. Image Process. | 4 |
| 2026 | AesPrompt: Zero-Shot Image Aesthetics Assessment With Multi-Granularity Aesthetic Prompt LearningabstractRecent years have witnessed increasing interest towards image aesthetics assessment (IAA), which predicts the aesthetic appeal of images by simulating human perception. The state-of-the-art IAA methods, despite their significant advancements, typically rely heavily on time-consuming and labor-intensive human annotation of aesthetic scores. Furthermore, they are subject to the generalization challenge, which is highly desired in real-world applications. Motivated by this, zero-shot image aesthetics assessment (ZIAA) is investigated to achieve robust model generalization without relying on manual aesthetic annotations, which remains largely underexplored. Specifically, a novel aesthetic prompt learning framework for ZIAA, dubbed AesPrompt, is presented in this paper. The key insight of AesPrompt is to emulate the human aesthetic perception process for learning aesthetic-oriented prompts in a multi-granularity manner. First, we first develop a new pseudo aesthetic distribution generation paradigm based on multi-LLM ensemble. Then, external knowledge of multi-granularity prompts encompassing image themes, emotions, and aesthetics is acquired. Through learning the multi-granularity aesthetic-oriented prompts, the proposed method achieves better generalization and interpretability. Extensive experiments on five IAA benchmarks demonstrate that AesPrompt consistently outperforms the state-of-the-art ZIAA methods across diverse-sourced images, covering natural images, artistic images, and artificial intelligence-generated images. Xiangfei Sheng, Leida Li, Pengfei Chen 0003, Giuseppe Valenzise |
IEEE Trans. Multim. | 5 |
| 2026 | Avatar Standardization Efforts for Interoperable Metaverse Services: Toward a Seamless Avatar-as-a-Service EcosystemabstractThe metaverse has emerged as a transformative platform that blends physical and digital realities, supporting immersive communication, interaction, and service delivery. At the core of these experiences are avatars, digital embodiments of users, that serve as interactive interfaces across a wide range of metaverse services, including virtual collaboration, education, healthcare, entertainment, e-commerce,etc.However, the fragmented ecosystem of tools, platforms, and proprietary formats has led to significant challenges in avatar interoperability, preventing seamless user identity and experience across applications. In response, numerous international standardization bodies have launched initiatives to define interoperable and modular avatar frameworks. This survey represents the first comprehensive review of these initiatives with a focus on service-oriented requirements, revealing the concept of Avatar-as-a-Service (AaaS). It examines both historical foundations and recent advances, with a particular attention given to the emerging ISO/IEC 23090-39 standard (MPEG-I Part 39 Avatar Representation Format, ARF). By synthesizing technical approaches, standardization efforts, and open challenges, this work presents a roadmap for the evolution of AaaS and highlights emerging opportunities for delivering personalized, high-fidelity, and cross-compatible avatar experiences to support the next generation of metaverse services. Anthony Trioux, Wei Zhang 0072, Giuseppe Valenzise, Fuzheng Yang 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | Blendshape Compression Techniques and Their Impact on Reconstructed Avatar Face Animation: A Subjective StudyabstractBlendshapes have been widely adopted as a key method for generating facial animation on avatars due to their ease of manipulation, flexibility in capturing diverse facial expressions, and compatibility with real-time rendering. However, current frameworks lack efficient methods for compressing blendshape (BS) animation parameters, which are critical for optimizing data transmission. This study introduces a pioneer compression scheme leveraging the amount of BS to be transmitted, their quantization as well as their transmission frequency. Subjective evaluation using the ITU-R BT.500-15 recommendation demonstrates that the proposed method significantly reduces the amount of data to transmit while preserving acceptable visual quality. This approach extends prior findings on reduced BS sets[kang2023effects] and addresses a significant gap in avatar media coding[avril2023morgan]. This work establishes a baseline for efficient facial animation and, serves as a foundation for further exploration on adaptive rate-allocation strategies and advanced compression strategies tailored for diverse avatar animation scenarios. Anthony Trioux, Wei Zhang 0072, Yusong Gao, Giuseppe Valenzise, Fuzheng Yang 0001 |
DCC | 5 |
| 2025 | Lift-PCAC: Lifting Based Point Cloud Attribute CompressionabstractPoint cloud (PC) compression is crucial for efficient transmission and storage in applications like virtual and augmented reality, where point counts can reach millions. While learning-based methods have shown promise in compressing PC geometry, attribute compression remains relatively unexplored. Existing methods often rely on variational autoencoders (VAEs). However, VAEs, with their low-dimensional bottlenecks, inherently limit the achievable reconstruction quality, especially at high bitrates. In this paper we introduce a novel approach to compressing PC attributes using a lifting framework. The invertibility of our method enables better reconstruction quality at high bitrates. Our Lifting Based Point Cloud Attribute Compression (Lift-PCAC) outperforms existing learning-based attribute compression methods on higher bitrates and shows comparable performance to G-PCC v.21 in some cases, highlighting the potential of this approach for point cloud compression. Rodrigo B. Pinheiro, Jean-Eudes Marvie, Giuseppe Valenzise, Frédéric Dufaux |
ICIP | 3 |
| 2025 | Denoising Diffusion Probabilistic Model for Point Cloud Compression at Low Bit-RatesabstractEfficient compression of low-bit-rate point clouds is critical for bandwidth-constrained applications. However, existing techniques mainly focus on high-fidelity reconstruction, requiring many bits for compression. This paper proposes a "Denoising Diffusion Probabilistic Model" (DDPM) architecture for point cloud compression (DDPM-PCC) at low bit-rates. A PointNet encoder produces the condition vector for the generation, which is then quantized via a learnable vector quantizer. This configuration allows to achieve a low bitrates while preserving quality. Experiments on ShapeNet and ModelNet40 show improved rate-distortion at low rates compared to standardized and state-of-the-art approaches. We publicly released the code at https://github.com/EIDOSLAB/DDPM-PCC. Gabriele Spadaro, Alberto Presta, Jhony-Heriberto Giraldo-Zuluaga, Marco Grangetto, Giuseppe Valenzise, Attilio Fiandrotti, Enzo Tartaglione |
ICME | 6 |
| 2025 | Non-Parametric Media Quality Recovery from Spammer-Affected Subjectively Annotated DatasetsabstractSpammer annotators in Quality of Experience (QoE) assessments provide unreliable ratings, often scoring randomly or assigning extreme ratings, which introduces noise and compromises data reliability. Modern methods to mitigate the effects of this noise rely on parametric approaches like maximum likelihood estimation or Bayesian techniques. These methods are sensitive to model assumptions and might therefore suffer from a lack of robustness. This paper proposes a non-parametric approach to measure annotator reliability and introduces the Non-Parametric subjective Quality Recovery (NPQR) algorithm, which is shown to compare favorably to state-of-the-art methods in terms of robustness against spammers. Lohic Fotio Tiotsop, Andrés Altieri, Giuseppe Valenzise |
ICME | 3 |
| 2025 | Exploring Compression Strategies for Blendshape-Based Avatar Facial Animation: Subjective and Objective AnalysisabstractBlendshapes (BS) have been widely adopted as a key method for generating facial animations on avatars due to their ease of manipulation, flexibility in capturing diverse expressions, and compatibility with real-time rendering. Despite their importance, current frameworks, including the recent MPEG Avatar Representation Format one, lack efficient methods for blendshape compression. This paper addresses this gap by proposing and analyzing a pioneering compression scheme for blendshape animation parameters. The method explores three key dimensions of compression: the number of transmitted BS, their quantization, as well as their frequency of transmission. The use of a linear interpolation for reduced blendshape transmission rates allow to mitigate flickering introduced by strong quantization, enhancing the viewing experience. Subjective evaluations demonstrate that the proposed approach achieves substantial data transmission savings while maintaining acceptable visual quality. Furthermore, an in-depth comparison of classical objective metrics against Mean Opinion Scores (MOS) reveals their limitations in accurately capturing perceived quality after blendshape compression, with Pearson and Spearman correlation scores reaching at most ~ 0.6. Among the evaluated metrics, Detail Loss Metric (DLM), Video Multimethod Assessment Fusion (VMAF), and Mean Peak Signal-to-Noise Ratio in the BS domain (PSNR-BS) exhibit the highest correlation with MOS. This study provides a comprehensive benchmark for blendshape compression and has driven the creation of an Exploratory Experiment (EE) within the ongoing MPEG avatar-related efforts, highlighting the study’s relevance to standardization activities. Anthony Trioux, Wei Zhang 0072, Giuseppe Valenzise, Fuzheng Yang 0001 |
ICME | 3 |
| 2025 | Improved RAHT-based Compression of 3D Gaussian Splats
Annalisa Gallina, Giuseppe Valenzise, Sara Baldoni, Federica Battisti |
PCS | 2 |
| 2024 | Reducing the Complexity of Normalizing Flow Architectures for Point Cloud Attribute CompressionabstractExisting learning-based methods to compress PCs attributes typically employ variational autoencoders (VAE) to learn compact signal representations. However, these schemes suffer from limited reconstruction quality at high bitrates due to their intrinsic lossy nature. More recently, normalizing flows (NF) have been proposed as an alternative solution. NFs are invertible networks that can achieve lossless reconstruction, at the cost of very large architectures with high memory and computational footprint. This paper proposes an improved NF architecture with reduced complexity called RNF-PCAC. It is composed of two operating modes specialized for low and high bitrates, combined in a rate-distortion optimized fashion. Our approach reduces the number of parameters of the existing NF architectures by over 6×. At the same time, it achieves state-of-the-art coding gains compared to previous learning-based methods and, for some PCs, it matches the performance of G-PCC (v.21). Rodrigo B. Pinheiro, Jean-Eudes Marvie, Giuseppe Valenzise, Frédéric Dufaux |
ICASSP | 3 |
| 2024 | Balancing Representation Abstractions and Local Details Preservation for 3d Point Cloud Quality Assessmentabstract3D Point Clouds (PCs) have become a valuable tool for representing intricate 3D information. Assessing the quality of PCs remains a challenging task, especially when striving for optimal immersive experiences. This paper introduces a novel metric and training approach that leverages projection-based views to evaluate the quality of 3D content. Our approach addresses a critical issue related to the intrinsic bias of deep networks for image recognition towards building hierarchical representations including only the global semantic, at the expense of local details. This bias is a limiting factor in tasks like 3D point cloud quality assessment where instances of the same content with varying degrees and types of degradation can possess strikingly similar representations. We propose a novel point cloud quality metric using a dual supervised and unsupervised training strategy to balance semantic understanding and preservation of critical perceptual quality-relevant information. The results demonstrate the effectiveness and reliability of our solution compared to state-of-the-art metrics on two standard 3D PCs quality assessment benchmarks (3D PCQA). Marouane Tliba, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux |
ICASSP | 3 |
| 2024 | A Toolkit to Benchmark Point Cloud Quality Metrics with Multi-Track Evaluation CriteriaabstractPoint clouds (PCs) gained popularity as a representation for 3D objects and scenes and are widely used in numerous applications in augmented and virtual reality domains. Concurrently, quality assessment of PCs became even more relevant to improve various aspects of these imaging pipelines. To stimulate further growth and interest in point cloud quality assessment (PCQA), we created a large-scale PCQA dataset (called “BASICS”) which provides the research community with a relevant and challenging dataset to develop reliable objective quality metrics, and we organized the PCVQA grand challenge at ICIP 2023. In this paper, we provide a track-based evaluation methodology for benchmarking visual quality metrics, mirroring the PCVQA grand challenge evaluation scenarios designed to mimic real-life applications. Furthermore, we provide a state-of-the-art benchmark for the point cloud quality metrics. The track-based benchmarking approach shows that there is room for improvement in certain research directions, drawing attention to open problems in the PCQA domain. Ali Ak, Emin Zerman, Maurice Quach, Aladine Chetouani, Giuseppe Valenzise, Patrick Le Callet |
ICIP | 5 |
| 2024 | Multi-Reference Generative Face Video Compression with Contrastive LearningabstractGenerative face video coding (GFVC) has been demonstrated as a potential approach to low-latency, low bitrate video conferencing. GFVC frameworks achieve an extreme gain in coding efficiency with over 70% bitrate savings when compared to conventional codecs at bitrates below 10kbps. In recent MPEG/JVET standardization efforts, all the information required to reconstruct video sequences using GFVC frameworks are adopted as part of the supplemental enhancement information (SEI) in existing compression pipelines. In light of this development, we aim to address a challenge that has been weakly addressed in prior GFVC frameworks, i.e., reconstruction drift as the distance between the reference and target frames increases. This challenge creates the need to update the reference buffer more frequently by transmitting more Intra-refresh frames, which are the most expensive element of the GFVC bitstream. To overcome this problem, we propose instead multiple reference animation as a robust approach to minimizing reconstruction drift, especially when used in a bi-directional prediction mode. Further, we propose a contrastive learning formulation for multi-reference animation. We observe that using a contrastive learning framework enhances the representation capabilities of the animation generator. The resulting framework, MRDAC (Multi-Reference Deep Animation Codec) can therefore be used to compress longer sequences with fewer reference frames or achieve a significant gain in reconstruction accuracy at comparable bitrates to previous frameworks. Quantitative and qualitative results show significant coding and reconstruction quality gains compared to previous GFVC methods, and more accurate animation quality in presence of large pose and facial expression changes. The source code will be available at https://github.com/Goluck-Konuko/animation-based-codecs Goluck Konuko, Giuseppe Valenzise |
MMSP | 2 |
| 2024 | Enhancing Immersive Experiences through 3D Point Cloud Analysis: A Novel Framework for Applying 2D Visual Saliency Models to 3D Point CloudsabstractIn the new area of immersive multimedia environments, understanding and manipulating visual attention are crucial for enhancing user experience. This study introduces an innovative framework that extends traditional 2D saliency maps to the analysis of 3D point clouds, a step forward in adapting saliency prediction to more complex and immersive environments. Our framework centers on the orthographic projection of 3D point clouds onto 2D planes, enabling the application of established 2D saliency models to this novel context. We further delve into the evaluation of these models on a 3D point cloud eye-tracking dataset, exploring various projection settings and thresholding techniques to maintain the integrity of saliency information in the transition from 2D to 3D. This research not only bridges a gap in applying visual attention models to 3D data but also offers insights into the optimization of quality of experience in immersive multimedia systems. Marouane Tliba, Xuemei Zhou, Irene Viola 0001, Pablo César, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux |
QoMEX | 6 |
| 2024 | Characterizing the Geometric Complexity of G-PCC Compressed Point CloudsabstractMeasuring the complexity of visual content is crucial in various applications, such as selecting sources to test processing algorithms, designing subjective studies, and efficiently determining the appropriate encoding parameters and bandwidth allocation for streaming. While spatial and temporal complexity measures exist for 2D videos, a geometric complexity measure for 3D content is still lacking. In this paper, we present the first study to characterize the geometric complexity of 3D point clouds. Inspired by existing complexity measures, we propose several compression-based definitions of geometric complexity derived from the rate-distortion curves obtained by compressing a dataset of point clouds using G-PCC. Additionally, we introduce density-based and geometry-based descriptors to predict complexity. Our initial results show that even simple density measures can accurately predict the geometric complexity of point clouds. Annalisa Gallina, Hadi Amirpour, Sara Baldoni, Giuseppe Valenzise, Federica Battisti |
VCIP | 4 |
| 2024 | Improving Reconstruction Fidelity in Generative Face Video Coding using High-Frequency ShuttlingabstractGenerative face video coding (GFVC) schemes applied to talking head videos have recently demonstrated significant coding gains compared to traditional coding frameworks, particularly at ultra-low bitrates. Despite advancements in the field, these methods still face challenges in handling large pose and facial expression changes, as well as (dis-)occlusions. Recently, a hybrid approach (HDAC+) that combines a low-quality video coded with a conventional codec and animation-based coding has been proposed for standardization and shown to partially address these issues. Although HDAC+ shows promising results, it still struggles with generating accurate images. In this paper, we propose HDAC-HF, an improvement to the reconstruction process in HDAC+. Based on empirical observations that some pose and expression details are lost during animation, we introduce a high-frequency (HF) shuttling mechanism to enhance reconstruction fidelity at the decoder side, inspired by recent advancements in video super-resolution. By enhancing the flow of high-frequency details in the feature domain, we improve the reconstruction of facial expressions and poses. Qualitative and quantitative experiments confirm that the proposed method improves reconstruction without any additional bitstream or signaling cost compared to the baseline HDAC+ codec. Goluck Konuko, Giuseppe Valenzise, Anthony Trioux |
VCIP | 2 |
| 2024 | BASICS: Broad Quality Assessment of Static Point Clouds in a Compression ScenarioabstractPoint clouds have become increasingly prevalent in representing 3D scenes within virtual environments, alongside 3D meshes. Their ease of capture has facilitated a wide array of applications on mobile devices, from smartphones to autonomous vehicles. Notably, point cloud compression has reached an advanced stage and has been standardized. However, the availability of quality assessment datasets, which are essential for developing improved objective quality metrics, remains limited. In this paper, we introduce BASICS, a large-scale quality assessment dataset tailored for static point clouds. The BASICS dataset comprises 75 unique point clouds, each compressed with four different algorithms including a learning-based method, resulting in the evaluation of nearly 1500 point clouds by 3500 unique participants. Furthermore, we conduct a comprehensive analysis of the gathered data, benchmark existing point cloud quality assessment metrics and identify their limitations. By publicly releasing the BASICS dataset, we lay the foundation for addressing these limitations and fostering the development of more precise quality metrics. Ali Ak, Emin Zerman, Maurice Quach, Aladine Chetouani, Aljoscha Smolic, Giuseppe Valenzise, Patrick Le Callet |
IEEE Trans. Multim. | 6 |
| 2024 | Subjective Media Quality Recovery From Noisy Raw Opinion Scores: A Non-Parametric PerspectiveabstractThis paper focuses on the challenge of accurately estimating the subjective quality of multimedia content from noisy opinion scores gathered from end-users. State-of-the-art methods rely on parametric statistical models to capture the subject's scoring behavior and recover quality estimates. However, these approaches have limitations, as they often require restrictive assumptions to achieve numerical stability during parameter estimation, leading to a lack of robustness when the modeling hypotheses do not fit the data. To overcome these limitations, we propose a paradigm shift towards non-parametric statistical methods. Specifically, we introduce a threefold contribution: i) in contrast to the prevailing approach in subjective quality recovery assuming a parametric score distribution, we propose a non parametric approach that guarantees greater accuracy by measuring reliability per subject and per stimulus, overcoming the limits of existing approaches that measure only per subject reliability; ii) we propose ESQR, a non-parametric algorithm for subjective quality recovery, demonstrating experimentally that it has higher robustness to noise compared to numerous state-of-the-art algorithms, thanks to the weaker assumptions made on data compared to parametric approaches; iii) the proposed approach is theoretically grounded, i.e., we define a non-parametric statistic and prove mathematically that it provides a measure of score reliability. The code to run ESQR and reproduce the results in this paper is made freely available at:http://media.polito.it/ESQR. Andrés Altieri, Lohic Fotio Tiotsop, Giuseppe Valenzise |
IEEE Trans. Multim. | 3 |
| 2023 | NF-PCAC: Normalizing Flow Based Point Cloud Attribute CompressionabstractLearning-based point cloud (PC) compression is a promising research avenue to reduce the transmission and storage costs for PC applications. Existing learning-based methods to compress PCs have mainly focused on geometry and employ variational autoencoders to learn compact signal representations. However, autoencoders leverage low-dimensional bottlenecks that limit the maximum reconstruction quality, even at high bitrates. In this paper, we propose a different and novel approach to compress PC attributes by using normalizing flows. Since normalizing flows model invertible transforms, the proposed approach can achieve better reconstruction quality than variational autoencoders over a large range of bitrates. Our Normalizing Flow-based Point Cloud Attribute Compression (NF-PCAC) outperforms previous learning-based methods for attribute compression, and has comparable performance as G-PCC v.14, showing the potential of this scheme for PC compression. Rodrigo B. Pinheiro, Jean-Eudes Marvie, Giuseppe Valenzise, Frédéric Dufaux |
ICASSP | 3 |
| 2023 | PCQA-Graphpoint: Efficient Deep-Based Graph Metric for Point Cloud Quality AssessmentabstractFollowing the advent of immersive technologies and the increasing interest in representing interactive geometrical format, 3D Point Clouds (PC) have emerged as a promising solution and effective means to display 3D visual information. In addition to other challenges in immersive applications, objective and subjective quality assessments of compressed 3D content remain open problems and an area of research interest. Yet most of the efforts in the research area ignore the local geometrical structures between points representation. In this paper, we overcome this limitation by introducing a novel and efficient objective metric for Point Clouds Quality Assessment, by learning local intrinsic dependencies using Graph Neural Network (GNN). To evaluate the performance of our method, two well-known datasets have been used. The results demonstrate the effectiveness and reliability of our solution compared to state-of-the-art metrics. Marouane Tliba, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux |
ICASSP | 3 |
| 2023 | Predictive Coding for Animation-Based Video CompressionabstractWe address the problem of efficiently compressing video for conferencing-type applications. We build on recent approaches based on image animation, which can achieve good reconstruction quality at very low bitrate by representing face motions with a compact set of sparse keypoints. However, these methods encode video in a frame-by-frame fashion, i.e., each frame is reconstructed from a reference frame, which limits the reconstruction quality when the bandwidth is larger. Instead, we propose a predictive coding scheme which uses image animation as a predictor, and codes the residual with respect to the actual target frame. The residuals can be in turn coded in a predictive manner, thus removing efficiently temporal dependencies. Our experiments indicate a significant bitrate gain, in excess of 70% compared to the HEVC video standard and over 30% compared to VVC, on a dataset of talking-head videos. Goluck Konuko, Stéphane Lathuilière, Giuseppe Valenzise |
ICIP | 3 |
| 2023 | Unified Measures for the Rate-Distortion-Latency Trade-offabstractIn today’s digital age, multimedia content is omnipresent, and the demand for efficient compression techniques is ever-increasing. In particular, the successful delivery of services based on video transmission largely depends on achieving the lowest latency values. One solution has been to use extrapolation for latency compensation in video transmission that allows to reduce the latency by an arbitrary amount. Nevertheless, this latency reduction comes at the cost of an increased distortion of the displayed images, since they are based on temporal extrapolation. Latency can also be traded with coding rate. This paper introduces ELR-PSNR and EPR-Latency as unified metrics to assess the three-way trade-off between rate, distortion, and latency simultaneously. Melan Vijayaratnam, Marta Milovanovic, Marco Cagnazzo, Enzo Tartaglione, Giuseppe Valenzise |
VCIP | 5 |
| 2022 | A Hybrid Deep Animation Codec for Low-Bitrate Video ConferencingabstractDeep generative models, and particularly facial animation schemes, can be used in video conferencing applications to efficiently compress a video through a sparse set of key-points, without the need to transmit dense motion vectors. While these schemes bring significant coding gains over con-ventional video codecs at low bitrates, their performance saturates quickly when the available bandwidth increases. In this paper, we propose a layered, hybrid coding scheme to overcome this limitation. Specifically, we extend a codec based on facial animation by adding an auxiliary stream con-sisting of a very low bitrate version of the video, obtained through a conventional video codec (e.g., HEVC). The an-imated and auxiliary videos are combined through a novel fusion module. Our results show consistent average BD-Rate gains in excess of -30% on a large dataset of video confer-encing sequences, extending the operational range of bitrates of a facial animation codec alone. Our code is available at github.com/animation-based-codecs Goluck Konuko, Stéphane Lathuilière, Giuseppe Valenzise |
ICIP | 3 |
| 2022 | Ultra-Low Bitrate Video Conferencing Using Deep Image AnimationabstractIn this work we propose a novel deep learning approach for ultra-low bitrate video compression for video conferencing applications. To address the shortcomings of current video compression paradigms when the available bandwidth is extremely limited, we adopt a model-based approach that employs deep neural networks to encode motion information as keypoint displacement and reconstruct the video signal at the decoder side. The overall system is trained in an end-to-end fashion minimizing a reconstruction error on the encoder output. Objective and subjective quality evaluation experiments demonstrate that the proposed approach provides an average bitrate reduction for the same visual quality of more than 60% compared to HEVC. Goluck Konuko, Giuseppe Valenzise, Stéphane Lathuilière |
ICIP | 2 |
| 2022 | Representation Learning Optimization for 3D Point Cloud Quality Assessment Without ReferenceabstractRecent information and communication systems have employed 3D Point Cloud (PC) as an advanced geometrical representation modality for immersive applications. Like most multimedia data, PCs are often compressed for transmission and viewing purposes, which can impact the perceived quality. Developing robust and efficient objective quality metrics for PCs is still an open problem. In this paper, we propose an end-to-end deep approach for evaluating the perceptual effects of point cloud compression solutions without reference. Our approach focuses on leveraging the intrinsic point cloud characteristics to quantify the coding impairments from few distant randomly selected patches using supervised and unsupervised training strategies. To evaluate the performance of our method, two well-known datasets have been used. The results demonstrate the effectiveness and reliability of the proposed method compared to to state-of-the-art methods. Marouane Tliba, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux |
ICIP | 3 |
| 2022 | Towards Zero-Latency Video Transmission Through Frame ExtrapolationabstractIn the past few years, several efforts have been devoted to reduce individual sources of latency in video delivery, including acquisition, coding and network transmission. The goal is to improve the quality of experience in applications requiring real-time interaction. Nevertheless, these efforts are fundamentally constrained by technological and physical limits. In this paper, we investigate a radically different approach that can arbitrarily reduce the overall latency by means of video extrapolation. We propose two latency compensation schemes where video extrapolation is performed either at the encoder or at the decoder side. Since a loss of fidelity is the price to pay for compensating latency arbitrarily, we study the latency-fidelity compromise using three recent video prediction schemes. Our preliminary results show that by accepting a quality loss, we can compensate a typical latency of 100 ms with a loss of 8 dB in PSNR with the best extrapolator. This approach is promising but also suggests that further work should be done in video prediction to pursue zero-latency video transmission. Melan Vijayaratnam, Marco Cagnazzo, Giuseppe Valenzise, Anthony Trioux, Michel Kieffer |
ICIP | 3 |
| 2021 | Ultra-Low Bitrate Video Conferencing Using Deep Image AnimationabstractIn this work we propose a novel deep learning approach for ultra-low bitrate video compression for video conferencing applications. To address the shortcomings of current video compression paradigms when the available bandwidth is extremely limited, we adopt a model-based approach that employs deep neural networks to encode motion information as keypoint displacement and reconstruct the video signal at the decoder side. The overall system is trained in an end-to-end fashion minimizing a reconstruction error on the encoder output. Objective and subjective quality evaluation experiments demonstrate that the proposed approach provides an average bitrate reduction for the same visual quality of more than 80% compared to HEVC. Goluck Konuko, Giuseppe Valenzise, Stéphane Lathuilière |
ICASSP | 2 |
| 2021 | Learning-Based Lossless Compression of 3D Point Cloud GeometryabstractThis paper presents a learning-based, lossless compression method for static point cloud geometry, based on context-adaptive arithmetic coding. Unlike most existing methods working in the octree domain, our encoder operates in a hybrid mode, mixing octree and voxel-based coding. We adaptively partition the point cloud into multi-resolution voxel blocks according to the point cloud structure, and use octree to signal the partitioning. On the one hand, octree representation can eliminate the sparsity in the point cloud. On the other hand, in the voxel domain, convolutions can be naturally expressed, and geometric information (i.e., planes, surfaces, etc.) is explicitly processed by a neural network. Our context model benefits from these properties and learns a probability distribution of the voxels using a deep convolutional neural network with masked filters, called VoxelDNN. Experiments show that our method outperforms the state-of-the-art MPEG G-PCC standard with average rate savings of 28% on a diverse set of point clouds from the Microsoft Voxelized Upper Bodies (MVUB) and MPEG. The implementation is available at https://github.com/Weafre/VoxelDNN. Dat Thanh Nguyen, Maurice Quach, Giuseppe Valenzise, Pierre Duhamel |
ICASSP | 3 |
| 2021 | Convolutional Neural Network for 3D Point Cloud Quality Assessment with ReferenceabstractIn recent years, the production of 3D content in the form of point clouds (PC) has increased considerably, especially in virtual reality applications. This enthusiasm is linked in particular to the development of acquisition technologies. In order to ensure a good quality of user experience, it is necessary to offer a high quality of visualization whatever the transmission medium used or the treatments applied. Thus, several metrics have been proposed which are essentially point-based metrics. In this article, we propose a deep learning-based method that efficiently predicts the quality of distorted PCs thanks to a set of features extracted from selected patches of the reference PC and its degraded version as well as the use of Convolutional Neural Networks (CNNs). The patches are selected randomly and the difference between corresponding patches is characterized by three attributes: geometry, curvature and color. The proposed method was evaluated and compared to state-of-the-art metrics using two datasets, including a large dataset more suited to deep learning models. We also compared different symmetrization functions and machine learning pooling as well as the ability of our method to predict the quality of unknown PCs through a cross-dataset evaluation. The results obtained show the relevance of the proposed framework with interesting perspectives. Aladine Chetouani, Maurice Quach, Giuseppe Valenzise, Frédéric Dufaux |
MMSP | 3 |
| 2021 | Learning-based lossless light field compressionabstractWe propose a learning-based method for lossless light field compression. The approach consists of two steps: first, the view to be compressed is synthesized based on previously decoded views; then, the synthesized view is used as a context to predict probabilities of the residual signal for adaptive arithmetic coding. We leverage recent advances in deep-learning-based view synthesis and generative modeling. Specifically, we evaluate two strategies for entropy modeling: a fully parallel probability estimation, where all pixel probabilities are estimated simultaneously; and a partially auto-regressive estimation, in which groups of pixels are predicted sequentially. Our results show that the latter approach provides the best coding gains compared to the state of the art, while keeping the computational complexity competitive. Milan Stepanov, M. Umair Mukati, Giuseppe Valenzise, Søren Forchhammer, Frédéric Dufaux |
MMSP | 3 |
| 2021 | A Perceptual Study of the Decoding Process of the SoftCast Wireless Video Broadcast SchemeabstractThe SoftCast scheme has been proposed as a promising alternative to traditional video broadcasting systems in wireless environments. In its current form, SoftCast performs image decoding at the receiver side by using a Linear Least Square Error (LLSE) estimator. Such approach maximizes the reconstructed quality in terms of Peak Signal-to-Noise Ratio (PSNR). However, we show that the LLSE induces an annoying blur effect at low Channel Signal-to-Noise Ratio (CSNR) quality. To cancel this artifact, we propose to replace the LLSE estimator by the Zero-Forcing (ZF) one. In order to better understand the perceived quality offered by these two estimators, a mathematical characterization as well as an objective and subjective studies are performed. Results show that the gains brought by the LLSE estimator, in terms of PSNR and Structural SIMiliraty (SSIM), are limited and quickly tend to null value as the CSNR increases. However, higher gains are obtained by the ZF estimator when considering the recent Video Multi-method Assessment Fusion (VMAF) metric proposed by Netflix, which evaluates the perceptual video quality. This result is confirmed by the subjective assessment. Anthony Trioux, Giuseppe Valenzise, Marco Cagnazzo, Michel Kieffer, François-Xavier Coudoux, Patrick Corlay, Mohamed Gharbi |
MMSP | 2 |
| 2021 | Lossless Coding of Point Cloud Geometry Using a Deep Generative ModelabstractThis paper proposes a lossless point cloud (PC) geometry compression method that uses neural networks to estimate the probability distribution of voxel occupancy. First, to take into account the PC sparsity, our method adaptively partitions a point cloud into multiple voxel block sizes. This partitioning is signalled via an octree. Second, we employ a deep auto-regressive generative model to estimate the occupancy probability of each voxel given the previously encoded ones. We then employ the estimated probabilities to code efficiently a block using a context-based arithmetic coder. Our context has variable size and can expand beyond the current block to learn more accurate probabilities. We also consider using data augmentation techniques to increase the generalization capability of the learned probability models, in particular in the presence of noise and lower-density point clouds. Experimental evaluation, performed on a variety of point clouds from four different datasets and with diverse characteristics, demonstrates that our method reduces significantly (by up to 37%) the rate for lossless coding compared to the state-of-the-art MPEG codec. Dat Thanh Nguyen, Maurice Quach, Giuseppe Valenzise, Pierre Duhamel |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Folding-Based Compression Of Point Cloud AttributesabstractExisting techniques to compress point cloud attributes leverage either geometric or video-based compression tools. We explore a radically different approach inspired by recent advances in point cloud representation learning. Point clouds can be interpreted as 2D manifolds in 3D space. Specifically, we fold a 2D grid onto a point cloud and we map attributes from the point cloud onto the folded 2D grid using a novel optimized mapping method. This mapping results in an image, which opens a way to apply existing image processing techniques on point cloud attributes. However, as this mapping process is lossy in nature, we propose several strategies to refine it so that attributes can be mapped to the 2D grid with minimal distortion. Moreover, this approach can be flexibly applied to point cloud patches in order to better adapt to local geometric complexity. In this work, we consider point cloud attribute compression; thus, we compress this image with a conventional 2D image codec. Our preliminary results show that the proposed folding-based coding scheme can already reach performance similar to the latest MPEG Geometry-based PCC (G-PCC) codec. Maurice Quach, Giuseppe Valenzise, Frédéric Dufaux |
ICIP | 2 |
| 2020 | Just Noticeable Quantization Levels For High Dynamic Range ImagesabstractJust noticeable quantization levels, which are conventionally used in picture coding, have been mainly developed for standard 8-bit images and low dynamic range (LDR) typical screens. The quantization levels however have not been adapted yet for high dynamic range (HDR) imaging and its accompanied HDR displays, which can reach up to a peak luminance of 4000 cd/m2. This study proposes an experimental methodology on HDR displays to determine just noticeable quantization levels for discrete cosine transform (DCT) coefficients on high luminance images. In the first stage of the proposed method, the quantization noise patterns for different DCT frequencies at different mean luminances are rendered by predicting the LED and LCD values of the two layer HDR display. Then, a two alternative forced choice based psychovisual experimental procedure using geometric search and QUEST methodology is realized by randomly presenting the rendered quantization noise at different amplitudes to the subjects in order to determine the just noticeable levels. The experiments are performed over 3 subjects for 30 different frequencies of 8×8 DCT patterns at mean luminances of 100 cd/m2and 1000 cd/m2. The results are interpreted with respect to frequency and luminance changes and from the point of utilized methodology, namely geometric search and QUEST. Sevim Begüm Sözer, Alper Koz, Ahmet Oguz Akyüz, Emin Zerman, Giuseppe Valenzise, Frédéric Dufaux |
ICIP | 5 |
| 2020 | Hybrid Learning-Based And Hevc-Based Coding Of Light FieldsabstractLight fields have additional storage requirements compared to conventional image and video signals, and demand therefore an efficient representation. In order to improve coding efficiency, in this work we propose a hybrid coding scheme which combines a learning-based compression approach with a traditional video coding scheme. Their integration offers great gains at low/mid bitrates thanks to the efficient representation of the learning-based approach and is competitive at high bitrates compared to standard tools thanks to the encoding of the residual signal. The proposed approach achieves on average 38% and 31% BD rate saving compared to HEVC and JPEG Pleno transform-based codec, respectively. Milan Stepanov, Giuseppe Valenzise, Frédéric Dufaux |
ICIP | 2 |
| 2020 | Improved Deep Point Cloud Geometry CompressionabstractPoint clouds have been recognized as a crucial data structure for 3D content and are essential in a number of applications such as virtual and mixed reality, autonomous driving, cultural heritage, etc. In this paper, we propose a set of contributions to improve deep point cloud compression, i.e.: using a scale hyperprior model for entropy coding; employing deeper transforms; a different balancing weight in the focal loss; optimal thresholding for decoding; and sequential model training. In addition, we present an extensive ablation study on the impact of each of these factors, in order to provide a better understanding about why they improve RD performance. An optimal combination of the proposed improvements achieves BD-PSNR gains over G-PCC trisoup and octree of 5.50 (6.48) dB and 6.84 (5.95) dB, respectively, when using the point-to-point (point-to-plane) metric. Code is available at https://github.com/mauriceqch/pcc_geo_cnn_v2. Maurice Quach, Giuseppe Valenzise, Frédéric Dufaux |
MMSP | 2 |
| 2020 | Subjective and Objective Quality Assessment of the SoftCast Video Transmission SchemeabstractSoftCast-based linear video coding and transmission (LVCT) schemes have been proposed as a promising alternative to traditional video coding and transmission schemes in wireless environments. Currently, the performance of LVCT schemes is evaluated by means of traditional objective scores such as PSNR or SSIM. Nevertheless, since the compression is performed in a very different way from traditional coding schemes such as HEVC, visual artifacts are also quite different and deserve to be subjectively assessed. In this paper, we propose a subjective quality assessment of SoftCast, pioneer and standard of the LVCT schemes. This study aims to better understand the trade-offs between the LVCT parameters that can be tuned to improve the quality. These parameters, including different GoP-sizes, Compression Ratios (CR) and Channel Signal-to-Noise Ratio (CSNR), are used to generate a dataset of 85 videos. A Double Stimulus Impairment Scale (DSIS) test is performed on the received videos to assess the perceived quality. Results show that the key characteristic of SoftCast, the linear relation between CSNR and PSNR, is also observed with the Mean-Opinion Scores (MOS), except at high CSNR where the quality saturates. In addition, Bjøntegaard model is used to quantify the trade-offs between CR, GoP-size and CSNR, depending on the intended application. Finally, the performance of objective metrics compared to the obtained MOS is evaluated. Results show that Multi-Scale SSIM (MS-SSIM), SSIM and Video Multimethod Assessment Fusion (VMAF) metrics offer the best correlation with the MOS values. Anthony Trioux, Giuseppe Valenzise, Marco Cagnazzo, Michel Kieffer, François-Xavier Coudoux, Patrick Corlay, Mohamed Gharbi |
VCIP | 2 |
| 2020 | From Pairwise Comparisons and Rating to a Unified Quality ScaleabstractThe goal of psychometric scaling is the quantification of perceptual experiences, understanding the relationship between an external stimulus, the internal representation and the response. In this paper, we propose a probabilistic framework to fuse the outcome of different psychophysical experimental protocols, namely rating and pairwise comparisons experiments. Such a method can be used for merging existing datasets of subjective nature and for experiments in which both measurements are collected. We analyze and compare the outcomes of both types of experimental protocols in terms of time and accuracy in a set of simulations and experiments with benchmark and real-world image quality assessment datasets, showing the necessity of scaling and the advantages of each protocol and mixing. Although most of our examples focus on image quality assessment, our findings generalize to any other subjective quality-of-experience task. María Pérez-Ortiz 0001, Aliaksei Mikhailiuk, Emin Zerman, Vedad Hulusic, Giuseppe Valenzise, Rafal Mantiuk |
IEEE Trans. Image Process. | 5 |
| 2020 | Deep Tone Mapping Operator for High Dynamic Range ImagesabstractA computationally fast tone mapping operator (TMO) that can quickly adapt to a wide spectrum of high dynamic range (HDR) content is quintessential for visualization on varied low dynamic range (LDR) output devices such as movie screens or standard displays. Existing TMOs can successfully tone-map only a limited number of HDR content and require an extensive parameter tuning to yield the best subjective-quality tone-mapped output. In this paper, we address this problem by proposing a fast, parameter-free and scene-adaptable deep tone mapping operator (DeepTMO) that yields a high-resolution and high-subjective quality tone mapped output. Based on conditional generative adversarial network (cGAN), DeepTMO not only learns to adapt to vast scenic-content (e.g., outdoor, indoor, human, structures, etc.) but also tackles the HDR related scene-specific challenges such as contrast and brightness, while preserving the fine-grained details. We explore 4 possible combinations of Generator-Discriminator architectural designs to specifically address some prominent issues in HDR related deep-learning frameworks like blurring, tiling patterns and saturation artifacts. By exploring different influences of scales, loss-functions and normalization layers under a cGAN setting, we conclude with adopting a multi-scale model for our task. To further leverage on the large-scale availability of unlabeled HDR data, we train our network by generating targets using an objective HDR quality metric, namely Tone Mapping Image Quality Index (TMQI). We demonstrate results both quantitatively and qualitatively, and showcase that our DeepTMO generates high-resolution, high-quality output images over a large spectrum of real-world scenes. Finally, we evaluate the perceived quality of our results by conducting a pair-wise subjective study which confirms the versatility of our method. Aakanksha Rana, Praveer Singh, Giuseppe Valenzise, Frédéric Dufaux, Nikos Komodakis, Aljoscha Smolic |
IEEE Trans. Image Process. | 3 |
| 2019 | Enhancing HEVC Spatial Prediction by Context-based LearningabstractDeep generative models have been recently employed to compress images, image residuals or to predict image regions. Based on the observation that state-of-the-art spatial prediction is highly optimized from a rate-distortion point of view, in this work we study how learning-based approaches might be used to further enhance this prediction. To this end, we propose an encoder-decoder convolutional network able to reduce the energy of the residuals of HEVC intra prediction, by leveraging the available context of previously decoded neigh-boring blocks. The proposed context-based prediction enhancement (CBPE) scheme enables to reduce the mean square error of HEVC prediction by 25% on average, without any additional signalling cost in the bitstream. Attilio Fiandrotti, Andrei I. Purica, Giuseppe Valenzise, Marco Cagnazzo |
ICASSP | 4 |
| 2019 | Learning Convolutional Transforms for Lossy Point Cloud Geometry CompressionabstractEfficient point cloud compression is fundamental to enable the deployment of virtual and mixed reality applications, since the number of points to code can range in the order of millions. In this paper, we present a novel data-driven geometry compression method for static point clouds based on learned convolutional transforms and uniform quantization. We perform joint optimization of both rate and distortion using a trade-off parameter. In addition, we cast the decoding process as a binary classification of the point cloud occupancy map. Our method outperforms the MPEG reference solution in terms of rate-distortion on the Microsoft Voxelized Upper Bodies dataset with 51.5% BDBR savings on average. Moreover, while octree-based methods face exponential diminution of the number of points at low bitrates, our method still produces high resolution outputs even at low bitrates. Code and supplementary material are available at https://github.com/mauriceqch/pcc_geo_cnn. Maurice Quach, Giuseppe Valenzise, Frédéric Dufaux |
ICIP | 2 |
| 2019 | Predicting Subjectivity in Image Aesthetics AssessmentabstractConventional image aesthetic quality prediction aims at predicting the average score of a picture or its aesthetic class (good/bad quality). However, aesthetic prediction is intrinsically subjective, and images with similar mean aesthetic scores/class might display very different levels of consensus by human raters. Recent work has dealt with aesthetic subjectivity by predicting the distribution of human scores. However, predicting the distribution is not directly interpretable in terms of subjectivity, and might be sub-optimal compared to directly estimating subjectivity descriptors computed from ground-truth scores. In this paper, we propose several measures of subjectivity, ranging from simple statistical measures such as the standard deviation of the scores, to newly proposed descriptors inspired by information theory. We evaluate the prediction performance of these measures when they are computed from predicted score distributions or when they are directly learned from ground-truth data. We find that the latter strategy provides in general better results, though there is still a large space for improvement in aesthetic subjectivity prediction. Chen Kang, Giuseppe Valenzise, Frédéric Dufaux |
MMSP | 2 |
| 2019 | Analysing the Impact of Cross-Content Pairs on Pairwise Comparison ScalingabstractPairwise comparisons (PWC) methodology is one of the most commonly used methods for subjective quality assessment, especially for computer graphics and multimedia applications. Unlike rating methods, a psychometric scaling operation is required to convert PWC results to numerical subjective quality values. Due to the nature of this scaling operation, the obtained quality scores are relative to the set they are computed in. While it is customary to compare different versions of the same content, in this work we study how cross-content comparisons may benefit psychometric scaling. For this purpose, we use two different video quality databases which have both rating and PWC experiment results. The results show that despite same-content comparisons play a major role in the accuracy of psychometric scaling, the use of a small portion of cross-content comparison pairs is indeed beneficial to obtain more accurate quality estimates. Emin Zerman, Giuseppe Valenzise, Aljoscha Smolic |
QoMEX | 2 |
| 2019 | An Adaptive Quantizer for High Dynamic Range Content: Application to Video CodingabstractIn this paper, we propose an adaptive perceptual quantization method to convert the representation of high dynamic range (HDR) content from the floating point data type to integer, which is compatible with the current image/video coding and display systems. The proposed method considers the luminance distribution of the HDR content, as well as the detectable contrast threshold of the human visual system, in order to preserve more contrast information than the perceptual quantizer (PQ) in integer representation. Aiming to demonstrate the effectiveness of this quantizer for HDR video compression, we implemented it in a mapping function on the top of the HDR video coding system based on high efficiency video coding standard. Moreover, a comparison function is also introduced to decrease the additional bit-rate of side information, generated by the mapping function. Objective quality measurements and subjective tests have been conducted in order to evaluate the quality of the reconstructed HDR videos. Subjective test results have shown that the proposed method can improve, in a significant manner, the perceived quality of some reconstructed HDR videos. In the objective assessment, the proposed method achieves improvements over PQ in terms of the average bit-rate gain for metrics used in the measurement. Yi Liu 0004, Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Giuseppe Valenzise, Emin Zerman |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Learning-Based Tone Mapping Operator for Efficient Image MatchingabstractIn this paper, we propose a new framework to optimally tone map the high dynamic range (HDR) content for image matching under drastic illumination variations. Since tone mapping operators (TMO) have traditionally been used for displaying HDR scenes, their design is suboptimal when used for computer vision tasks, such as image matching. We address this suboptimality by proposing a two-step framework, consisting of: first, a luminance-invariant guidance model based on a support vector regressor (SVR) to optimally adapt the tone mapping function for image matching; and second, an energy maximization model to generate appropriate training samples for learning the SVR. At each step, we collectively address both stages of keypoint detection and descriptor extraction in the feature matching framework. By locally altering the intrinsic characteristics of the tone mapping function, the learned guidance model facilitates the extraction of local invariant features in the presence of illumination variations. We demonstrate that the proposed TMO significantly outperforms perceptually driven state-of-the-art TMOs on a dataset of HDR scenes characterized by challenging lighting variations, such as day/night transitions. Aakanksha Rana, Giuseppe Valenzise, Frédéric Dufaux |
IEEE Trans. Multim. | 2 |
| 2018 | Learning Local Distortion Visibility from Image QualityabstractAccurate prediction of local distortion visibility thresholds is critical in many image and video processing applications. Existing methods require an accurate modeling of the human visual system, and are derived through pshycophysical experiments with simple, artificial stimuli. These approaches, however, are difficult to generalize to natural images with complex types of distortion. In this paper, we explore a different perspective, and we investigate whether it is possible to learn local distortion visibility from image quality scores. We propose a convolutional neural network based optimization framework to infer local detection thresholds in a distorted image. Our model is trained on multiple quality datasets, and the results are correlated with empirical visibility thresholds collected on complex stimuli in a recent study. Our results are comparable to state-of-the-art mathematical models that were trained on phsycovisual data directly. This suggests that it is possible to predict psychophysical phenomena from visibility information embedded in image quality scores. Navaneeth K. Kottayil, Irene Cheng 0001, Giuseppe Valenzise, Frédéric Dufaux |
ICIP | 3 |
| 2018 | Quality Assessment of Deep-Learning-Based Image CompressionabstractImage compression standards rely on predictive coding, transform coding, quantization and entropy coding, in order to achieve high compression performance. Very recently, deep generative models have been used to optimize or replace some of these operations, with very promising results. However, so far no systematic and independent study of the coding performance of these algorithms has been carried out. In this paper, for the first time, we conduct a subjective evaluation of two recent deep-learning-based image compression algorithms, comparing them to JPEG 2000 and to the recent BPG image codec based on HEVC Intra. We found that compression approaches based on deep auto-encoders can achieve coding performance higher than JPEG 2000, and sometimes as good as BPG. We also show experimentally that the PSNR metric is to be avoided when evaluating the visual quality of deep-learning-based methods, as their artifacts have different characteristics from those of DCT or wavelet-based codecs. In particular, images compressed at low bitrate appear more natural than JPEG 2000 coded pictures, according to a no-reference naturalness measure. Our study indicates that deep generative models are likely to bring huge innovation into the video coding arena in the coming years. Giuseppe Valenzise, Andrei I. Purica, Vedad Hulusic, Marco Cagnazzo |
MMSP | 1 |
| 2018 | Video Quality Evaluation for Tile-Based Spatial AdaptationabstractThe following topics are dealt with: learning (artificial intelligence); feature extraction; video coding; object detection; data compression; image classification; image coding; image representation; image reconstruction; optimisation. Hiba Yousef, Jean Le Feuvre, Giuseppe Valenzise, Vedad Hulusic |
MMSP | 3 |
| 2018 | TRISK: A local features extraction framework for texture-plus-depth content matching
Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux |
Image Vis. Comput. | 2 |
| 2018 | Spatio-temporal constrained tone mapping operator for HDR video compression
Cagri Ozcinar, Paul Lauga, Giuseppe Valenzise, Frédéric Dufaux |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Fine-grained detection of inverse tone mapping in HDR images
Wei Fan 0004, Giuseppe Valenzise, Francesco Banterle, Frédéric Dufaux |
Signal Process. | 2 |
| 2018 | Blind Quality Estimation by Disentangling Perceptual and Noisy Features in High Dynamic Range ImagesabstractHigh dynamic range (HDR) image visual quality assessment in the absence of a reference image is challenging. This research topic has not been adequately studied largely due to the high cost of HDR display devices. Nevertheless, HDR imaging technology has attracted increasing attention, because it provides more realistic content, consistent to what the human visual system perceives. We propose a new no-reference image quality assessment (NR-IQA) model for HDR data based on convolutional neural networks. The proposed model is able to detect visual artifacts, taking into consideration perceptual masking effects, in a distorted HDR image without any reference. The error and perceptual masking values are measured separately, yet sequentially, and then processed by a mixing function to predict the perceived quality of the distorted image. Instead of using simple stimuli and psychovisual experiments, perceptual masking effects are computed from a set of annotated HDR images during our training process. Experimental results demonstrate that our proposed NR-IQA model can predict HDR image quality as accurately as state-of-the-art full-reference IQA methods. Navaneeth K. Kottayil, Giuseppe Valenzise, Frédéric Dufaux, Irene Cheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Good features to track for RGBD imagesabstractRGBD (texture-plus-depth) image representation enriches traditional 2D content with additional geometrical information, having the potential to improve the performance of many computer vision tasks. In image matching, this has been partially studied by considering how depth maps can help render feature descriptors more distinctive. However, little has been done to design keypoint detection approaches able to leverage the availability of depth information. In this paper, we propose a novel and robust approach for detecting corners from RGBD images. Our method modifies a classical corner detection strategy, based on local second-order moment matrices, by computing derivatives in a coordinate system which reflects the local properties of object surfaces. Our results demonstrate a higher stability to out-of-plane rotations of the proposed RGBD corner detector both in terms of feature repeatability and in a visual odometry application. Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux |
ICASSP | 2 |
| 2017 | An adaptive perceptual quantization method for HDR video codingabstractThis paper presents a new adaptive perceptual quantization method for the High Dynamic Range (HDR) content. This method considers the luminance distribution of the HDR image as well as the Minimum Detectable Contrast (MDC) thresholds to preserve the contrast information during quantization. Base on this method, we develop a mapping function for HDR video compression and apply it to a HEVC Main 10 Profile-based video coding chain. Our experiments show that the proposed mapping function can efficiently improve the quality of the reconstructed HDR video in both objective and subjective assessments. Yi Liu 0004, Naty Ould Sidaty, Wassim Hamidouche, Olivier Déforges, Giuseppe Valenzise, Emin Zerman |
ICIP | 5 |
| 2017 | Learning-based tone mapping operator for image matchingabstractIn this paper, we propose a new framework to optimally tone-map a high dynamic range (HDR) content for image matching under drastic illumination variations. This task is of fundamental importance for many computer vision applications. To design such a framework, we build a luminance invariant guidance model using a Support Vector Regressor (SVR) and learn it to facilitate the extraction of invariant descriptors from scenes subject to wide variety of appearance changes such as day/night transition. To this end, we initially generate appropriate training samples using a simple similarity-maximization mechanism. We then employ the learned model to predict optimal modulation maps that help to locally alter the intrinsic characteristics (such as shape, size) of the tone mapping function. We evaluate the proposed model performance in terms of matching score and mean average precision rate using state-of-the-art descriptor extraction schemes. We demonstrate that our tone mapping framework significantly outperforms the existing perceptually-driven state-of-the-art TMOs on the benchmark datasets. Aakanksha Rana, Giuseppe Valenzise, Frédéric Dufaux |
ICIP | 2 |
| 2017 | Learning-based adaptive tone mapping for keypoint detectionabstractThe goal of tone mapping operators (TMOs) has traditionally been to display high dynamic range (HDR) pictures in a perceptually favorable way. However, when tone-mapped images are to be used for computer vision tasks such as keypoint detection, these design approaches are suboptimal. In this paper, we propose a new learning-based adaptive tone mapping framework which aims at enhancing keypoint stability under drastic illumination variations. To this end, we design a pixel-wise adaptive TMO which is modulated based on a model derived by Support Vector Regression (SVR) using local higher order characteristics. To circumvent the difficulty to train SVR in this context, we further propose a simple detection-similarity-maximization model to generate appropriate training samples using multiple images undergoing illumination transformations. We evaluate the performance of our proposed framework in terms of keypoint repeatability for state-of-the-art keypoint detectors. Experimental results show that our proposed learning-based adaptive TMO yields higher keypoint stability when compared to existing perceptually-driven state-of-the-art TMOs. Aakanksha Rana, Giuseppe Valenzise, Frédéric Dufaux |
ICME | 2 |
| 2017 | Statistical analysis and directional coding of layer-based HDR image coding residueabstractExisting methods for layer-based backward compatible high dynamic range (HDR) image and video coding mostly focus on the rate-distortion optimization of base layer while neglecting the encoding of the residue signal in the enhancement layer. Although some recent studies handle residue coding by designing function based fixed global mapping curves for 8-bit conversion and exploiting standard codecs on the resulting 8-bit images, they do not take the local characteristics of residue blocks into account. Inspired by the local anisotropic characteristics of the residue signal and directional methods for motion compensated low dynamic range (LDR) video coding, in this paper we first investigate whether HDR image coding residue exhibits also local anisotropic characteristics. Specifically, we verify directional structures in residue blocks by means of auto-covariance analysis for different bitrates, spatial activities and dynamic ranges as the main variables in HDR image coding. Then, we compare the rate distortion performances of directional coding methods with the baseline residue coding methods in the literature along with different combinations of 8-bit conversion methods. The experiments indicate that content dependent 8-bit conversions and directional coding significantly outperforms the existing function based 8-bit conversions and typical coding for residue coding. Kutan Feyiz, Fatih Kamisli, Emin Zerman, Giuseppe Valenzise, Alper Koz, Frédéric Dufaux |
MMSP | 4 |
| 2017 | Quality of experience in UHD-1 phase 2 television: The contribution of UHD+HFR technologyabstractA key factor to determine the quality of experience (QoE) of a video is its capability to convey the large spectrum of perceptual phenomena that our eyes can sense in real life. In order to meet this demand, the recent DVB UHD-1 Phase 2 specification employs new video features, such as higher spatial resolutions (4K/8K) and High Frame Rate (HFR). The first enables larger field of view and level of details, while the second offers sharper images of moving objects going well beyond the current frame rates. While the contribution of each of these technologies to QoE has been investigated individually, in this paper we are interested to study their interaction, and in quantifying the benefits to users from their combination. To this end, we conduct a subjective test on compressed UHD+HFR content on a recent display capable of reproducing 100 pictures per second at 2160p resolution, with the goal to assess the increase in QoE of UHD and HFR with respect to conventional video, both individually and in combination. The results indicate that for content with fast motion, at higher bitrates the combination of UHD and HFR significantly improves the QoE compared to that obtained when these features are used individually. Vedad Hulusic, Giuseppe Valenzise, Jean-Charles Gicquel, Jérôme Fournier, Frédéric Dufaux |
MMSP | 2 |
| 2017 | Effect of color space on high dynamic range video compression performanceabstractHigh dynamic range (HDR) technology allows for capturing and delivering a greater range of luminance levels compared to traditional video using standard dynamic range (SDR). At the same time, it has brought multiple challenges in content distribution, one of them being video compression. While there has been a significant amount of work conducted on this topic, there are some aspects that could still benefit this area. One such aspect is the choice of color space used for coding. In this paper, we evaluate through a subjective study how the performance of HDR video compression is affected by three color spaces: the commonly used Y'CbCr, and the recently introduced ITP (ICtCp) and Ypu'v'. Five video sequences are compressed at four bit rates, selected in a preliminary study, and their quality is assessed using pairwise comparisons. The results of pairwise comparisons are further analyzed and scaled to obtain quality scores. We found no evidence of ITP improving compression performance over Y'CbCr. We also found that Ypu'v' results in a moderately lower performance for some sequences. Emin Zerman, Vedad Hulusic, Giuseppe Valenzise, Rafal Mantiuk, Frédéric Dufaux |
QoMEX | 3 |
| 2017 | A model of perceived dynamic range for HDR images
Vedad Hulusic, Kurt Debattista, Giuseppe Valenzise, Frédéric Dufaux |
Signal Process. Image Commun. | 3 |
| 2016 | An image smoothing operator for fast and accurate scale space approximationabstractGussian image smoothing is a fundamental operation in the extraction of scale-invariant feature points. Its computation, however, can be too expensive in some resource-constrained scenarios. Alternative solutions such as the box filter can be computed more efficiently, at the cost of a loss in feature repeatibility under some conditions. In this paper we propose a fast and accurate image smoothing operator based on integral images. It has the same order of computational complexity as the box filter, but provides much more accurate visual results and improved keypoint repeatability, which is confirmed in a feature detection scenario using SIFT features. Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux |
ICASSP | 2 |
| 2016 | Forensic detection of inverse tone mapping in HDR imagesabstractHigh dynamic range (HDR) imaging is attracting an increasing deal of attention in the multimedia community, yet its forensic problems have been little studied so far. This paper proposes an HDR image forensic method, which aims at differentiating HDR images created from multiple low dynamic range (LDR) images from those created from a single LDR image by inverse tone mapping. For each kind of HDR image, a Gaussian mixture model is learned. Thereafter, an HDR image forensic feature is constructed based on calculating the Fisher scores. With comparison to a steganalytic feature and a texture/facial analysis feature, experimental results demonstrate the efficiency of the proposed method in HDR image forensic classification on whole images as well as small blocks, for three inverse tone mapping methods. Wei Fan 0004, Giuseppe Valenzise, Francesco Banterle, Frédéric Dufaux |
ICIP | 2 |
| 2016 | Optimizing tone mapping operators for keypoint detection under illumination changesabstractTone mapping operators (TMO) have recently raised interest for their capability to handle illumination changes. However, these TMOs are optimized with respect to perception rather than image analysis tasks like key point detection. Moreover, no work has been done to analyze the factors affecting the optimization of TMOs for such tasks. In this paper, we investigate the influence of two factors-Correlation Coefficient (CC) and Repeatability Rate (RR) of the tone mapped images for the optimization of classical Retinex based models to enhance key point detection under illumination changes. CC-based optimized models aim at increasing the similarity of the tone mapped images. Conversely, RR-based optimized models quantify the optimal detection performance gains. By considering two simple Retinex based models, i.e., Gaussian and bilateral filtering, we show that estimating as precisely as possible the illumination, CC-based optimized models do not necessarily bring to optimal key point detection performance. We conclude that, instead, other criteria specific to RR-based optimized models should be taken into account. Moreover, large gains in performance with respect to existing popular TMOs motivate further research towards optimal tone mapping technique for computer vision applications. Aakanksha Rana, Giuseppe Valenzise, Frédéric Dufaux |
MMSP | 2 |
| 2016 | Perceived dynamic range of HDR imagesabstractAlthough high dynamic range (HDR) imaging has gained great popularity and acceptance in both the scientific and commercial domains, the relationship between perceptually accurate, content-independent dynamic range and objective measures has not been fully explored. In this paper, a new methodology for perceived dynamic range evaluation of complex stimuli in HDR conditions is proposed. A subjective study with 20 participants was conducted and correlations between mean opinion scores (MOS) and three image features were analyzed. Strong Spearman correlations between MOS and objective DR measure and between MOS and image key were found. An exploratory analysis reveals that additional image characteristics should be considered when modeling perceptually-based dynamic range metrics. Finally, one of the outcomes of the study is the perceptually annotated HDR image dataset with MOS values, that can be used for HDR imaging algorithms and metric validation, content selection and analysis of aesthetic image attributes. Vedad Hulusic, Giuseppe Valenzise, Edoardo Provenzi, Kurt Debattista, Frédéric Dufaux |
QoMEX | 2 |
| 2016 | Using region-of-interest for quality evaluation of DIBR-based view synthesis methodsabstractAs 3D media became more and more popular over the last years, new technologies are needed in the transmission, compression and creation of 3D content. One of the most commonly used techniques for aiding with the compression and creation of 3D content is known as view synthesis. The most effective class of view synthesis algorithms are using Depth-Image-Based-Rendering techniques, which use explicit scene geometry to render new views. However, these methods may produce geometrical distortions and localized artifacts which are difficult to evaluate as they are inherently different from encoding errors and they are perceived differently by human subjects. In this paper, we propose a region-of-interest evaluation technique for view synthesis based on DIBR methods. Based on the assumption that certain areas determined by the geometrical properties of the scene are prone to distortions, we select a ROI by analyzing the multiple DIBR methods together with the ground truth. The approach is tested using a subjective evaluation view synthesis database and show that our method improves the SSIM correlation with subjective scores We also test another similar method and traditional metrics. Andrei I. Purica, Giuseppe Valenzise, Béatrice Pesquet-Popescu, Frédéric Dufaux |
QoMEX | 2 |
| 2016 | An evaluation of HDR image matching under extreme illumination changesabstractHigh dynamic range (HDR) imaging has potential to facilitate computer vision tasks such as image matching where lighting transformations hinder the matching performance. However, little has been done to quantify the gains with different possible HDR representations for vision algorithms like feature extraction. In this paper, we evaluate the performance of the full feature extraction pipeline, including detection and description, on ten different image representations: low dynamic range (LDR), seven different tone mapped (TM) HDR and two HDR imaging (linear and log encoded) representations. We measure the impact of using these different representations for feature matching using mean average precision (mAP) scores on four illumination change datasets. We perform feature extraction using four popular schemes in the literature: SIFT, SURF, BRISK, FREAK. With respect to previous studies, our observations confirm the advantages of HDR over conventional LDR imagery, and the fact that HDR linear values are not appropriate for vision tasks. However, HDR representations that work best for keypoint detection are not necessarily optimal when the full feature extraction is taken into account. Aakanksha Rana, Giuseppe Valenzise, Frédéric Dufaux |
VCIP | 2 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 27 |
| 2016 | Keypoint Detection in RGBD Images Based on an Anisotropic Scale SpaceabstractThe increasing availability of texture+depth (RGBD) content has recently motivated research toward the design of image features able to employ the additional geometrical information provided by depth. Indeed, such features are supposed to provide higher robustness than conventional 2D features in the presence of large changes of camera viewpoint. In this paper, we consider the first stage of RGBD image matching, i.e., keypoint detection. In order to obtain viewpoint-covariant keypoints, we design a filtering process, which approximates a diffusion process along the surfaces of the scene, by means of the information provided by depth. Next, we employ this multiscale representation to find keypoints through a multiscale keypoint detector. The keypoints obtained by the proposed detector provide substantially higher stability to viewpoint changes than alternative 2D and RGBD feature extraction approaches, both in terms of repeatability and image classification accuracy. Furthermore, the proposed detector can be efficiently implemented on a GPU. Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux |
IEEE Trans. Multim. | 2 |
| 2015 | Improving distinctiveness of brisk features using depth mapsabstractBinary local descriptors are widely used in computer vision thanks to their compactness and robustness to many image transformations such as rotations or scale changes. However, more complex transformations, like changes in camera viewpoint, are difficult to deal with using conventional features due to the lack of geometric information about the scene. In this paper, we propose a local binary descriptor which assumes that geometric information is available as a depth map. It employs a local parametrization of the scene surface, obtained through depth information, which is used to build a BRISK-like sampling pattern intrinsic to the scene surface. Although we illustrate the proposed method using the BRISK architecture, the obtained parametrization is rather general and could be embedded into other binary descriptors. Our simulations on a set of synthetically generated scenes show that the proposed descriptor is significantly more stable and distinctive than popular BRISK descriptors under a wide range of viewpoint angle changes. Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux |
ICIP | 2 |
| 2015 | A scale space for texture+depth images based on a discrete laplacian operatorabstractIn this paper we design a smoothing filter for texture+depth images based on anisotropic diffusion. Our proposed filter enables to generate a scale space on the texture image guided by depth information, and is linear and numerically stable. We show experimentally that using scene geometry preserves the internal structure of 3D surfaces (e.g., it avoids smoothing across object boundaries). As a consequence, the result of smoothing is more independent to changes in the camera position. To illustrate the practical utility of a scale space with such properties, we integrate our filter into the SIFT keypoint detector, getting a substantial improvement of the repeatability of detected keypoints under significant viewpoint position changes. Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux |
ICME | 2 |
| 2015 | Evaluation of Feature Detection in HDR Based Imaging Under Changes in Illumination ConditionsabstractHigh dynamic range (HDR) imaging enables to capture details in both dark and very bright regions of a scene, and is therefore supposed to provide higher robustness to illumination changes than conventional low dynamic range (LDR) imaging in tasks such as visual features extraction. However, it is not clear how much this gain is, and which are the best modalities of using HDR to obtain it. In this paper we evaluate the first block of the visual feature extraction pipeline, i.e., keypoint detection, using both LDR and different HDR-based modalities, when significant illumination changes are present in the scene. To this end, we captured a dataset with two scenes and a wide range of illumination conditions. On these images, we measure how the repeatability of either corner or blob interest points is affected with different LDR/HDR approaches. Our observations confirm the potential of HDR over conventional LDR acquisition. Moreover, extracting features directly from HDR pixel values is more effective than first tonemapping and then extracting features, provided that HDR luminance information is previously encoded to perceptually linear values. Aakanksha Rana, Giuseppe Valenzise, Frédéric Dufaux |
ISM | 2 |
| 2014 | Local visual features extraction from texture+depth content based on depth image analysisabstractWith the increasing availability of low-cost - yet precise - depth cameras, “texture+depth” content has become more and more popular in several computer vision and 3D rendering tasks. Indeed, depth images bring enriched geometrical information about the scene which would be hard and often impossible to estimate from conventional texture pictures. In this paper, we investigate how the geometric information provided by depth data can be employed to improve the stability of local visual features under a large spectrum of viewpoint changes. Specifically, we leverage depth information to derive local projective transformations and compute descriptor patches from the texture image. Since the proposed approach may be used with any blob detector, it can be seamlessly integrated into the processing chain of state-of-the-art visual features such as SIFT. Our experiments show that a geometry-aware feature extraction can bring advantages in terms of descriptor distinctiveness with respect to state-of-the-art scale and affine-invariant approaches. Maxim Karpushin, Giuseppe Valenzise, Frédéric Dufaux |
ICIP | 2 |
| 2014 | Detectability-quality trade-off in JPEG counter-forensicsabstractRemoving JPEG quantization footprints from an image inevitably introduces artifacts and traces in the spatial domain. Recently, several robust methods have been proposed to detect footprints of counter-forensics and recover the image's compression history. In this paper we investigate the limitations of these detectors, by proposing an improved counter-forensic attack which adds a postprocessing denoising step besides dithering. We consider both a general-purpose denoising algorithm and one targeted to JPEG images. In the latter case, we show that this approach can successfully reduce the accuracy of detectors in the literature to that of a random decision. As a second contribution, we study the trade-off between the detectability of counter-forensics and quality of the tampered image, and show that the loss of quality is not sufficient for the analyst to use available no-reference quality assessment tools as an indicator of an attack. Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro |
ICIP | 1 |
| 2013 | Segmentation-based optimized tone mapping for high dynamic range image and video codingabstractA core part of the state-of-the art high dynamic range (HDR) image and video compression methods is the tone mapping operation to convert the visible luminance range into the finite bit depths that can be supported by the current video codecs. These conversions are until now optimized to provide backward compatibility to the existing low dynamic range (LDR) displays. However, a direct application of these methods for the emerging HDR displays can result in a loss of details in the bright and dark regions of the HDR content. In this paper, we overcome this limitation by designing a tone mapping operation which handles the bright and dark regions separately. The proposed method first finds the optimal segmentation of the HDR image into two parts, namely dark and bright regions, and then designs the optimal tone mapping for each region in terms of the mean square error between the logarithm of the luminance values of the original and reconstructed HDR content (HDR-MSE). The results indicate the superiority of the proposed method over the state-of-the art HDR coding methods. Paul Lauga, Alper Koz, Giuseppe Valenzise, Frédéric Dufaux |
PCS | 3 |
| 2013 | Revealing the Traces of JPEG Compression Anti-ForensicsabstractDue to the lossy nature of transform coding, JPEG introduces characteristic traces in the compressed images. A forensic analyst might reveal these traces by analyzing the histogram of discrete cosine transform (DCT) coefficients and exploit them to identify local tampering, copy-move forgery, etc. At the same time, it has been recently shown that a knowledgeable adversary can possibly conceal the traces of JPEG compression, by adding a dithering noise signal in the DCT domain, in order to restore the histogram of the original image. In this paper, we study the processing chain that arises in the case of JPEG compression anti-forensics. We take the perspective of the forensic analyst, and we show how it is possible to counter the aforementioned anti-forensic method revealing the traces of JPEG compression, regardless of the quantization matrix being used. Tests on a large image dataset demonstrated that the proposed detector was able to achieve an average accuracy equal to 93%, rising above 99% when excluding the case of nearly lossless JPEG compression. Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2012 | Motion prediction of depth video for depth-image-based rendering using don't care regionsabstractTo enable synthesis of any desired intermediate view between two captured views at decoder via depth-image-based rendering (DIBR), both texture and depth maps from the captured viewpoints must be encoded and transmitted in a format known as texture-plus-depth. In this paper, we focus on the compression of depth maps across time to lower the overall bitrate in texture-plus-depth format. We observe that depth maps are not directly viewed, but are only used to provide geometric information of the captured scene for view synthesis at decoder. Thus, as long as the resulting geometric error does not lead to unacceptable synthesized view quality, each depth pixel only needs to be reconstructed at the decoder coarsely within a tolerable range. We first formalize the notion of tolerable range per depth pixel as don't care region (DCR), by studying the synthesized view distortion sensitivity to the pixel value - a sensitive depth pixel will have a narrow DCR, and vice versa. Given per-pixel DCRs, we then modify inter-prediction modes during motion prediction to search for a predictor block matching per-pixel DCRs in a target block (rather than the fixed ground truth depth signal in a target block), in order to lower the energy of the prediction residual for the block. We implemented our DCR-based motion prediction scheme inside H.264; our encoded bitstreams remain 100% standard compliant. We show experimentally that our proposed encoding scheme can reduce the bitrate of depth maps coded with baseline H.264 by over 28%. Giuseppe Valenzise, Gene Cheung, Rafael Galvão de Oliveira, Marco Cagnazzo, Béatrice Pesquet-Popescu, Antonio Ortega |
PCS | 1 |
| 2012 | No-Reference Pixel Video Quality Monitoring of Channel-Induced DistortionabstractVideo transmitted over an error-prone network may be received at the decoder with degradations due to packet losses. No-reference quality monitoring algorithms are the most practical way to measure the quality of the received video, since they do not impose any change with respect to the network architecture. Conventionally, these methods assume the availability of the corrupted bitstream. In some situations this is not possible, e.g., because the bitstream is encrypted or processed by third-party decoders, and only the decoded pixel values can be used. The major issue in this scenario is the lack of knowledge about which regions of the video have been actually lost, which is a fundamental ingredient for estimating channel-induced distortion. In this paper, we propose a maximum a posteriori estimation of the pattern of lost macroblocks, which assumes the knowledge of the decoded pixels only. This information can be used as input to a no-reference quality monitoring system, which produces an accurate estimate of the mean-square-error (MSE) distortion introduced by channel errors. The results of the proposed method are well correlated with the MSE distortion computed in full-reference mode, with a linear correlation coefficient equal to 0.9 at frame level and 0.98 at sequence level. Giuseppe Valenzise, Stefano Magni, Marco Tagliasacchi, Stefano Tubaro |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | The cost of JPEG compression anti-forensicsabstractThe statistical footprint left by JPEG compression can be a valuable source of information for the forensic analyst. Recently, it has been shown that a suitable anti-forensic method can be used to destroy these traces, by properly adding a noise-like signal to the quantized DCT coefficients. In this paper we analyze the cost of this technique in terms of introduced distortion and loss of image quality. We characterize the dependency of the distortion on the image statistics in the DCT domain and on the quantization step used in JPEG compression. We also evaluate the loss of quality as measured by means of a perceptual metric, showing that a perceptually-optimized version of the anti-forensic method fails to completely conceal the forgery. Our conclusion is that removing the traces of the JPEG compression history could be much more challenging than it might appear, as anti-forensic methods are bound to leave characteristic traces. Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro |
ICASSP | 1 |
| 2011 | Countering JPEG anti-forensicsabstractJPEG coding leaves characteristic footprints that can be leveraged to reveal doctored images, e.g. providing the evidence for local tampering, copy-move forgery, etc. Recently, it has been shown that a knowledgeable attacker might attempt to remove such footprints by adding a suitable anti-forensic dithering signal to the image in the DCT domain. Such noise-like signal restores the distribution of the DCT coefficients of the original picture, at the cost of affecting image quality. In this paper we show that it is possible to detect this kind of attack by measuring the noisiness of images obtained by re-compressing the forged image at different quality factors. When tested on a large set of images, our method was able to correctly detect forged images in 97% of the cases. In addition, the original quality factor could be accurately estimated. Giuseppe Valenzise, Vitaliano Nobile, Marco Tagliasacchi, Stefano Tubaro |
ICIP | 1 |
| 2010 | A reduced-reference structural similarity approximation for videos corrupted by channel errors
Marco Tagliasacchi, Giuseppe Valenzise, Matteo Naccari, Stefano Tubaro |
Multim. Tools Appl. | 2 |
| 2010 | Joint Compressive Video Coding and AnalysisabstractTraditionally, video acquisition, coding and analysis have been designed and optimized as independent tasks. This has a negative impact in terms of consumed resources, as most of the raw information captured by conventional acquisition devices is discarded in the coding phase, while the analysis step only requires a few descriptors of salient video characteristics. Recent compressive sensing literature has partially broken this paradigm by proposing to integrate sensing and coding in a unified architecture composed by a light encoder and a more complex decoder, which exploits sparsity of the underlying signal for efficient recovery. However, a clear understanding of how to embed video analysis in this scheme is still missing. In this paper, we propose a joint compressive video coding and analysis scheme and, as a specific application example, we consider the problem of object tracking in video sequences. We show that, weaving together compressive sensing and the information computed by the analysis module, the bit-rate required to perform reconstruction and tracking of the foreground objects can be considerably reduced, with respect to a conventional disjoint approach that postpones the analysis after the video signal is recovered in the pixel domain. These findings suggest that considerable gains in performance can be potentially obtained in video analysis applications, provided that a joint analysis-aware design of acquisition, coding and signal recovery is carried out. M. Cossalter, Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro |
IEEE Trans. Multim. | 2 |
| 2009 | Privacy-Enabled Object Tracking in Video Sequences Using Compressive SensingabstractIn a typical video analysis framework, video sequences are decoded and reconstructed in the pixel domain before being processed for high level tasks such as classification or detection.Nevertheless, in some application scenarios, it might be of interest to complete these analysis tasks without disclosing sensitive data, e.g. the identity of people captured by surveillance cameras. In this paper we propose a new coding scheme suitable for video surveillance applications that allows tracking of video objects without the need to reconstruct the sequence,thus enabling privacy protection. By taking advantage of recent findings in the compressive sensing literature, we encode a video sequence with a limited number of pseudo-random projections of each frame. At the decoder, we exploit the sparsity that characterizes background subtracted images in order to recover the location of the foreground object. We also leverage the prior knowledge about the estimated location of the object, which is predicted by means of a particle filter, to improve the recovery of the foreground object location. The proposed framework enables privacy, in the sense it is impossible to reconstruct the original video content from the encoded random projections alone, as well as secrecy, since decoding is prevented if the seed used to generate the random projections is not available. M. Cossalter, Marco Tagliasacchi, Giuseppe Valenzise |
AVSS | 3 |
| 2009 | A reduced-reference video structural similarity metric based on no-reference estimation of channel-induced distortionabstractThe reduced-reference (RR) approximation of a full-reference (FR) video quality assessment method is a convenient way to build evaluation metrics which are both intrinsically well correlated with human judgments and feasible to implement in a network scenario, without the need to explore the perceptual significance of new video features through mean opinion score tests. In this paper, we propose a RR approximation of the video structural similarity index (VSSIM), a FR metric which is known to be well descriptive of the video quality perceived by users. We focus on the visual degradation produced by channel transmission errors: first, at the encoder, a small set of salient structural video features is assembled and transmitted through the RR channel to the end-user; then, at the decoder the feature vector is combined with a fine-granularity, no-reference estimate of the channel-induced distortion to produce the VSSIM approximation. By uniformly quantizing the feature vector and compressing it using a context-adaptive, variable length encoder, we show that good correlation coefficients with ground-truth VSSIM (rho = 0.85) may be achieved spending, respectively, less than 12 and 27 kbps for a video sequence with CIF or SD resolution. Andrea Albonico, Giuseppe Valenzise, Matteo Naccari, Marco Tagliasacchi, Stefano Tubaro |
ICASSP | 2 |
| 2009 | A compressive-sensing based watermarking scheme for sparse image tampering identificationabstractIn this paper we describe a robust watermarking scheme for image tampering identification and localization. A compact representation of the image is first produced by assembling a feature vector consisting of pseudo-random projections of the decimated image. Then, the quantized projections are encoded to form a hash, which is robustly embedded as a watermark in the image. By recovering the watermark the random projections are obtained, and then used to estimate the distortion of the received image. If tampering is sufficiently sparse or compressible in some basis description, a map of the introduced modification is recovered. The system relies on compressive sensing and distributed source coding principles to reduce the size of the hash of a 1024 × 1024 image, to about 4,000 bits. With this hash length, tampering sparse up to 20% and with a tampering energy around a PSNR of 15 dB can be successfully localized. Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro, Giacomo Cancelli, Mauro Barni |
ICIP | 1 |
| 2009 | Hash-Based Identification of Sparse Image TamperingabstractIn the last decade, the increased possibility to produce, edit, and disseminate multimedia contents has not been adequately balanced by similar advances in protecting these contents from unauthorized diffusion of forged copies. When the goal is to detect whether or not a digital content has been tampered with in order to alter its semantics, the use of multimedia hashes turns out to be an effective solution to offer proof of legitimacy and to possibly identify the introduced tampering. We propose an image hashing algorithm based on compressive sensing principles, which solves both the authentication and the tampering identification problems. The original content producer generates a hash using a small bit budget by quantizing a limited number of random projections of the authentic image. The content user receives the (possibly altered) image and uses the hash to estimate the mean square error distortion between the original and the received image. In addition, if the introduced tampering is sparse in some orthonormal basis or redundant dictionary, an approximation is given in the pixel domain. We emphasize that the hash is universal, e.g., the same hash signature can be used to detect and identify different types of tampering. At the cost of additional complexity at the decoder, the proposed algorithm is robust to moderate content-preserving transformations including cropping, scaling, and rotation. In addition, in order to keep the size of the hash small, hash encoding/decoding takes advantage of distributed source codes. Marco Tagliasacchi, Giuseppe Valenzise, Stefano Tubaro |
IEEE Trans. Image Process. | 2 |
| 2008 | Resource constrained efficient acoustic source localization and tracking using a distributed network of microphonesabstractIn this paper we present an efficient method to perform acoustic source localization and tracking using a distributed network of microphones. In this scenario, there is a trade-off between the localization performance and the expense of resources: in fact, a minimization of the localization error would require to use as many sensors as possible; at the same time, as the number of microphones increases, the cost of the network inevitably tends to grow, while in practical applications only a limited amount of resources is available. Therefore, at each time instant only a subset of the sensors should be enabled in order to meet the cost constraints. We propose a heuristic method for the optimal selection of this subset of microphones, using as distortion metrics the Cramer-Rao lower bound (CRLB) and as cost function the total distance between the selected sensors. The heuristic approach has been compared to an optimal algorithm, which searches the best sensor configuration among the full set of microphones, while satisfying the cost constraint. The proposed heuristic algorithm yields similar performance w.r.t. the full-search procedure, but at a much less computational cost. We show that this method can be used effectively in an acoustic source tracking application. Giuseppe Valenzise, Giorgio Prandi, Marco Tagliasacchi, Augusto Sarti |
ICASSP | 1 |
| 2008 | Minimum variance multiplexing of multimedia objectsabstractThis paper addresses the problem of simultaneous transmission of multiple multimedia objects (such as images or video sequences) over a bandwidth-limited channel. The trivial strategy of partitioning in equal parts the available rate among the bitstreams is suboptimal, when the multimedia objects have different coding complexities. Exploiting object diversity allows us to allocate the bandwidth according to some optimality criteria, e.g. minimizing the average total distortion or minimizing the variance between the distortions of each object. By describing the rate-distortion characteristics of each multimedia object in terms of a simple exponential model, we provide a closed form solution for both the minimum average and the minimum variance problems. In addition, if we consider the statistical distribution of the rate-distortion model parameters, we can show that the minimum variance solution can effectively reduce the quality fluctuations among the objects, with an overall coding efficiency loss, w.r.t. the minimum average solution, of only 0.5dB on average. Some experiments, carried out on different H.264/AVC video sequences, validate our theoretical results. Giuseppe Valenzise, Marco Tagliasacchi, Stefano Tubaro |
ICASSP | 1 |
| 2008 | Localization of sparse image tampering via random projectionsabstractHashes can be used to provide authentication of multimedia contents. In the case of images, a hash can be used to detect whether the data has been modified in an illegitimate way. When the authentication check fails, it might be useful to localize the tampering in the spatial domain. This paper proposes an algorithm based on compressive sensing principles, which solves both the authentication and the localization problems. The encoder produces a hash using a small bit budget by quantizing a limited number of random projections of the authentic image. The decoder uses the hash to estimate the distortion between the original and the received image. In addition, if the attack is sparse, it can be also localized. In order to keep the size of the hash small, encoding/decoding takes advantage of distributed source codes. This paper also investigates experimentally the tradeoff between the rate allocated to the hash and the performance achieved in terms of tampering localization. Marco Tagliasacchi, Giuseppe Valenzise, Stefano Tubaro |
ICIP | 2 |
| 2008 | Reduced-reference estimation of channel-induced video distortion using distributed source codingabstractChannel-induced distortion estimation is an important aspect in the delivery of video contents over IP networks: the QoS requirements of both content providers and content users conflict with the intrinsic best-effort nature of packet-switched networks, which may introduce annoying artifacts in the received streams due to channel errors or jitter. In this paper we propose a Reduced-Reference video quality assessment method, based on objective quality metrics, which enables distortion estimation at the macroblock level. The content provider transmits a small feature vector for each frame, starting from random projections computed for each macroblock. In order to reduce the bit rate of the transmitted feature vector, we encode it using Distributed Source Coding (DSC) tools. The content user decodes the feature vector using the received sequence as side information. Additionally, the end-user may take advantage of some prior information about the support of the errors in the frame in such a way that the required bit length of the transmitted feature vector is further reduced. In our experiments, using 4 random projections, the use of DSC enables a bit saving of 70% w.r.t. scalar quantization and transmission of the original feature vector; when also the a priori error map is available at the decoder, the average length of the transmitted partial reference can be further reduced by another 5% of average. Giuseppe Valenzise, Matteo Naccari, Marco Tagliasacchi, Stefano Tubaro |
ACM Multimedia | 1 |
| 2008 | Minimum Variance Optimal Rate Allocation for Multiplexed H.264/AVC BitstreamsabstractConsider the problem of transmitting multiple video streams to fulfill a constant bandwidth constraint. The available bit budget needs to be distributed across the sequences in order to meet some optimality criteria. For example, one might want to minimize the average distortion or, alternatively, minimize the distortion variance, in order to keep almost constant quality among the encoded sequences. By working in the rho-domain, we propose a low-delay rate allocation scheme that, at each time instant, provides a closed form solution for either the aforementioned problems. We show that minimizing the distortion variance instead of the average distortion leads, for each of the multiplexed sequences, to a coding penalty less than 0.5 dB, in terms of average PSNR. In addition, our analysis provides an explicit relationship between model parameters and this loss. In order to smooth the distortion also along time, we accommodate a shared encoder buffer to compensate for rate fluctuations. Although the proposed scheme is general, and it can be adopted for any video and image coding standard, we provide experimental evidence by transcoding bitstreams encoded using the state-of-the-art H.264/AVC standard. The results of our simulations reveal that is it possible to achieve distortion smoothing both in time and across the sequences, without sacrificing coding efficiency. Marco Tagliasacchi, Giuseppe Valenzise, Stefano Tubaro |
IEEE Trans. Image Process. | 2 |
| 2007 | Scream and gunshot detection and localization for audio-surveillance systemsabstractThis paper describes an audio-based video surveillance system which automatically detects anomalous audio events in a public square, such as screams or gunshots, and localizes the position of the acoustic source, in such a way that a video-camera is steered consequently. The system employs two parallel GMM classifiers for discriminating screams from noise and gunshots from noise, respectively. Each classifier is trained using different features, chosen from a set of both conventional and innovative audio features. The location of the acoustic source which has produced the sound event is estimated by computing the time difference of arrivals of the signal at a microphone array and using linear-correction least square localization algorithm. Experimental results show that our system can detect events with a precision of 93% at a false rejection rate of 5% when the SNR is 10dB, while the source direction can be estimated with a precision of one degree. A real-time implementation of the system is going to be installed in a public square of Milan. Giuseppe Valenzise, Luigi Gerosa, Marco Tagliasacchi, Fabio Antonacci, Augusto Sarti |
AVSS | 1 |