EDBT 2026 Demo / reviewers in the wild / expert
Marco Cagnazzo
dblp:13/2693
· DBLP profile ↗
88ranked-venue papers
18as first author
14since 2021 · last 2026
0000-0001-6731-3755ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 84 · 15 first-author · 13 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-authorComputer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Compression in 3D Gaussian Splatting: A Survey of Methods, Trends, and Future Directionsabstract3D Gaussian Splatting (3DGS) has recently emerged as a pioneering approach in explicit scene rendering and computer graphics. Unlike traditional neural radiance field (NeRF) methods, which typically rely on implicit, coordinate-based models to map spatial coordinates to pixel values, 3DGS utilizes millions of learnable 3D Gaussians. Its differentiable rendering technique and inherent capability for explicit scene representation and manipulation positions 3DGS as a potential game-changer for the next generation of 3D reconstruction and representation technologies. This enables 3DGS to deliver real-time rendering speeds while offering unparalleled editability levels. However, despite its advantages, 3DGS suffers from substantial memory and storage requirements, posing challenges for deployment on resource-constrained devices. In this survey, we provide a comprehensive overview focusing on the scalability and compression of 3DGS. We begin with a detailed background overview of 3DGS, followed by a structured taxonomy of existing compression methods. Additionally, we analyze and compare current methods from the topological perspective, evaluating their strengths and limitations in terms of fidelity, compression ratios, and computational efficiency. Furthermore, we explore how advancements in efficient NeRF representations can inspire future developments in 3DGS optimization. Finally, we conclude with current research challenges and highlight key directions for future exploration. Muhammad Salman Ali, Chaoning Zhang, Marco Cagnazzo, Giuseppe Valenzise, Enzo Tartaglione, Sung-Ho Bae |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Robust and efficient airplane cockpit video coding leveraging temporal redundancyabstractAbstract Airplane cockpit screens consist of virtual instruments where characters, numbers, and graphics are overlaid on a black or natural background. Recording the cockpit screen allows one to log vital plane data, as aircraft manufacturers do not offer direct access to raw data. However, traditional video codecs struggle at preserving character readability at the required low bit-rates. We showed in a previous work that large rate-distortion gains can be achieved if the characters are encoded as text rather than as pixels. We now leverage temporal redundancy to both achieve robust character recognition and improve encoding efficiency. A convolutional neural network is trained for character classification over synthetic samples augmented with occlusions to gain robustness against overlapping graphics. Further robustness to background occlusions is brought by a probabilistic framework that error-corrects the output of the convolutional neural network. Next, we propose a predictive text coding technique specifically tailored for text in cockpit videos that achieves competitive performance over commodity lossless methods. Experiments with real cockpit video footage show large rate-distortion gains for the proposed method with respect to three different video compression standards. Notably, the H.264/AVC codec retrofitted with our method outperforms H.265/HEVC-SCC and is competitive with the much more complex H.266/VVC while preserving text and graphics. The entire pipeline described in this work has been implemented at Safran Electronics as an embedded avionics system drawing just 2W of power thanks to a combination of software and FPGA implementation. Iulia Mitrica, Attilio Fiandrotti, Christophe Ruellan, Marco Cagnazzo |
Multim. Tools Appl. | 4 |
| 2024 | Find the Lady: Permutation and Re-synchronization of Deep Neural NetworksabstractDeep neural networks are characterized by multiple symmetrical, equi-loss solutions that are redundant. Thus, the order of neurons in a layer and feature maps can be given arbitrary permutations, without affecting (or minimally affecting) their output. If we shuffle these neurons, or if we apply to them some perturbations (like fine-tuning) can we put them back in the original order i.e. re-synchronize? Is there a possible corruption threat? Answering these questions is important for applications like neural network white-box watermarking for ownership tracking and integrity verification. We advance a method to re-synchronize the order of permuted neurons. Our method is also effective if neurons are further altered by parameter pruning, quantization, and fine-tuning, showing robustness to integrity attacks. Additionally, we provide theoretical and practical evidence for the usual means to corrupt the integrity of the model, resulting in a solution to counter it. We test our approach on popular computer vision datasets and models, and we illustrate the threat and our countermeasure on a popular white-box watermarking method. Carl De Sousa Trias, Mihai Mitrea, Attilio Fiandrotti, Marco Cagnazzo, Sumanta Chaudhuri, Enzo Tartaglione |
AAAI | 4 |
| 2024 | WaterMAS: Sharpness-Aware Maximization for Neural Network Watermarking
Carl De Sousa Trias, Mihai Mitrea, Attilio Fiandrotti, Marco Cagnazzo, Sumanta Chaudhuri, Enzo Tartaglione |
ICPR (5) | 4 |
| 2023 | Glass-to-Glass Delay Reduction: Encoding Rate Reduction vs. Video Frame ExtrapolationabstractApplications such as teleoperated driving, remote robot control, and telepresence rely on video services to ensure real-time interaction with a satisfying quality of experience. Reducing the Glass-to-Glass (G2G) delay, i.e., the time delay between the acquisition of a video frame and its display on a remote terminal is critical for these applications. Deep learning-based video frame extrapolation before video encoding has been recently considered as an interesting solution to reduce G2G delay, however, the latency introduced by extrapolation has not been taken into account. In this paper, considering the main sources of latency, including extrapolation delay, we examine the benefits and limitations of frame extrapolation at encoder in reducing the G2G delay in a point-to-point video transmission system. To this end, we compare the latency-quality trade-off for two latency compensation methods: encoding rate reduction and video frame extrapolation. Our aim is to determine the G2G delay reduction that may be achieved at the price of a given quality reduction. Our experiments show that extrapolation methods can provide a null perceived G2G delay with an acceptable loss in quality, particularly for applications with video contents with limited temporal information. Such delay reduction is unreachable via encoding rate reduction. Hind Kanj, Anthony Trioux, Marco Cagnazzo, François-Xavier Coudoux, Patrick Corlay, Michel Kieffer |
MMSP | 3 |
| 2023 | Unified Measures for the Rate-Distortion-Latency Trade-offabstractIn today’s digital age, multimedia content is omnipresent, and the demand for efficient compression techniques is ever-increasing. In particular, the successful delivery of services based on video transmission largely depends on achieving the lowest latency values. One solution has been to use extrapolation for latency compensation in video transmission that allows to reduce the latency by an arbitrary amount. Nevertheless, this latency reduction comes at the cost of an increased distortion of the displayed images, since they are based on temporal extrapolation. Latency can also be traded with coding rate. This paper introduces ELR-PSNR and EPR-Latency as unified metrics to assess the three-way trade-off between rate, distortion, and latency simultaneously. Melan Vijayaratnam, Marta Milovanovic, Marco Cagnazzo, Enzo Tartaglione, Giuseppe Valenzise |
VCIP | 3 |
| 2022 | Towards Zero-Latency Video Transmission Through Frame ExtrapolationabstractIn the past few years, several efforts have been devoted to reduce individual sources of latency in video delivery, including acquisition, coding and network transmission. The goal is to improve the quality of experience in applications requiring real-time interaction. Nevertheless, these efforts are fundamentally constrained by technological and physical limits. In this paper, we investigate a radically different approach that can arbitrarily reduce the overall latency by means of video extrapolation. We propose two latency compensation schemes where video extrapolation is performed either at the encoder or at the decoder side. Since a loss of fidelity is the price to pay for compensating latency arbitrarily, we study the latency-fidelity compromise using three recent video prediction schemes. Our preliminary results show that by accepting a quality loss, we can compensate a typical latency of 100 ms with a loss of 8 dB in PSNR with the best extrapolator. This approach is promising but also suggests that further work should be done in video prediction to pursue zero-latency video transmission. Melan Vijayaratnam, Marco Cagnazzo, Giuseppe Valenzise, Anthony Trioux, Michel Kieffer |
ICIP | 2 |
| 2022 | Depth Patch Selection for Decoder-Side Depth Estimation in MPEG Immersive VideoabstractThe MPEG immersive video (MIV) standard has been developed to efficiently compress volumetric video content and enable an immersive user experience. MIV deals with an enormous amount of data that comes in the form of multi-view plus depth videos, which is efficiently reduced in the process of pruning, by tackling the redundancies among the views. This paper presents a novel approach for improving the existing immersive video coding scheme. The proposed approach reduces the amount of transmitted depth data, leveraging the fact that the depth information is partially contained in texture videos. The study proposes a method that ensures a reliable recovery of depths at the decoder-side. This method provides BD-rate improvements on both high and low bitrate ranges, with up to 22.57% Y-PSNR, 25.76%VMAF, 24.07% MS-SSIM, and 22.94% IV-PSNR metric gain, given a low bitrate setting. Marta Milovanovic, Félix Henry, Marco Cagnazzo |
PCS | 3 |
| 2022 | Online Learning for Adaptive Video Streaming in Mobile NetworksabstractIn this paper, we propose a novel algorithm for video bitrate adaptation in HTTP Adaptive Streaming (HAS), based on online learning. The proposed algorithm, named Learn2Adapt (L2A) , is shown to provide a robust bitrate adaptation strategy which, unlike most of the state-of-the-art techniques, does not require parameter tuning, channel model assumptions, or application-specific adjustments. These properties make it very suitable for mobile users, who typically experience fast variations in channel characteristics. Experimental results, over real 4G traffic traces, show that L2A improves on the overall Quality of Experience (QoE) and in particular the average streaming bitrate, a result obtained independently of the channel and application scenarios. Theodoros Karagkioules, Georgios S. Paschos, Nikolaos Liakopoulos, Attilio Fiandrotti, Dimitrios Tsilimantos, Marco Cagnazzo |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2021 | Patch Decoder-Side Depth Estimation In Mpeg Immersive VideoabstractThis paper presents a new approach for achieving bitrate and pixel rate reduction in the MPEG immersive video coding setting. We demonstrate that it is possible to avoid the transmission of some depth information in the Test Model for Immersive Video (TMIV) by estimating it at the receiver's side. Although the transmitted information in TMIV is considered as non-redundant, we show that it is possible to improve this algorithm. This method provides 3.4%, 9.0%, and 12.1% average BD-rate gain for natural content on high, medium, and low bitrate, respectively, with up to respectively 12.3%, 16.0%, and 18.4% peak reductions. Moreover, it preserves the perceptual quality as measured with MS-SSIM and VMAF metrics. Additionally, it decreases the pixel rate by 8.3% for each test sequence. Marta Milovanovic, Félix Henry, Marco Cagnazzo, Joël Jung |
ICASSP | 3 |
| 2021 | A Perceptual Study of the Decoding Process of the SoftCast Wireless Video Broadcast SchemeabstractThe SoftCast scheme has been proposed as a promising alternative to traditional video broadcasting systems in wireless environments. In its current form, SoftCast performs image decoding at the receiver side by using a Linear Least Square Error (LLSE) estimator. Such approach maximizes the reconstructed quality in terms of Peak Signal-to-Noise Ratio (PSNR). However, we show that the LLSE induces an annoying blur effect at low Channel Signal-to-Noise Ratio (CSNR) quality. To cancel this artifact, we propose to replace the LLSE estimator by the Zero-Forcing (ZF) one. In order to better understand the perceived quality offered by these two estimators, a mathematical characterization as well as an objective and subjective studies are performed. Results show that the gains brought by the LLSE estimator, in terms of PSNR and Structural SIMiliraty (SSIM), are limited and quickly tend to null value as the CSNR increases. However, higher gains are obtained by the ZF estimator when considering the recent Video Multi-method Assessment Fusion (VMAF) metric proposed by Netflix, which evaluates the perceptual video quality. This result is confirmed by the subjective assessment. Anthony Trioux, Giuseppe Valenzise, Marco Cagnazzo, Michel Kieffer, François-Xavier Coudoux, Patrick Corlay, Mohamed Gharbi |
MMSP | 3 |
| 2021 | A Multi-View Stereoscopic Video Database With Green Screen (MTF) For Video Transition Quality-of-Experience AssessmentabstractWe introduce a multi-view stereoscopic video database with a green screen, called MTF, for the usages in computer vision applications, in particular for free navigation, free-viewpoint television, and video transition quality-of-experience (QoE) assessment. The MTF contains full-HD videos of real storytelling made up of 3 scenes. One particularity of this dataset is that to understand its storytelling, users must change their point of view in the scene at a given time. To this end, we usually need to generate a transition to link two points of view in the same scene. Computer vision techniques that enable such transitions like view synthesis methods, rely on a set of images of the scene to render some new views from different viewpoints of this scene. However, these methods may have many failure cases that lead to artifacts in the final rendered video transition. In most view synthesis QoE tests, the contents are not designed to make the transition between two points of view useful or interesting for the viewers, e.g. they don't need to make a transition to capture more information to better understand the content. We thus, assume that participants will harshly judge artifacts and imperfections in the rendered transition. Thus, the MTF is expected to enable a better analysis of the visual impact of persistent artifacts in the final rendered transition. In our dataset, all the scenes are recorded in a green screen studio, which is often used to superimpose special effects and scenery during editing according to specific needs. Our dataset also presents a wide baseline camera-setup, a challenging constraint for view synthesis techniques. Finally, The MTF can also be used as a complementary dataset with others in literature in various computer vision applications, such as video compression, 3D video content, immersive virtual reality environment, optical flow estimation... Nour Hobloss, Lu Zhang 0037, Marco Cagnazzo |
QoMEX | 3 |
| 2021 | HEMP: High-order entropy minimization for neural network compression
Enzo Tartaglione, Stéphane Lathuilière, Attilio Fiandrotti, Marco Cagnazzo, Marco Grangetto |
Neurocomputing | 4 |
| 2021 | Hybrid dual stream blender for wide baseline view synthesis
Nour Hobloss, Lu Zhang 0037, Stéphane Lathuilière, Marco Cagnazzo, Attilio Fiandrotti |
Signal Process. Image Commun. | 4 |
| 2020 | DR2S: Deep Regression with Region Selection for Camera Quality EvaluationabstractIn this work, we tackle the problem of estimating a camera capability to preserve fine texture details at a given lighting condition. Importantly, our texture preservation measurement should coincide with human perception. Consequently, we formulate our problem as a regression one and we introduce a deep convolutional network to estimate texture quality score. At training time, we use ground-truth quality scores provided by expert human annotators in order to obtain a subjective quality measure. In addition, we propose a region selection method to identify the image regions that are better suited at measuring perceptual quality. Finally, our experimental evaluation shows that our learning-based approach outperforms existing methods and that our region selection algorithm consistently improves the quality estimation. Marcelin Tworski, Stéphane Lathuilière, Salim Belkarfa, Attilio Fiandrotti, Marco Cagnazzo |
ICPR | 5 |
| 2020 | Subjective and Objective Quality Assessment of the SoftCast Video Transmission SchemeabstractSoftCast-based linear video coding and transmission (LVCT) schemes have been proposed as a promising alternative to traditional video coding and transmission schemes in wireless environments. Currently, the performance of LVCT schemes is evaluated by means of traditional objective scores such as PSNR or SSIM. Nevertheless, since the compression is performed in a very different way from traditional coding schemes such as HEVC, visual artifacts are also quite different and deserve to be subjectively assessed. In this paper, we propose a subjective quality assessment of SoftCast, pioneer and standard of the LVCT schemes. This study aims to better understand the trade-offs between the LVCT parameters that can be tuned to improve the quality. These parameters, including different GoP-sizes, Compression Ratios (CR) and Channel Signal-to-Noise Ratio (CSNR), are used to generate a dataset of 85 videos. A Double Stimulus Impairment Scale (DSIS) test is performed on the received videos to assess the perceived quality. Results show that the key characteristic of SoftCast, the linear relation between CSNR and PSNR, is also observed with the Mean-Opinion Scores (MOS), except at high CSNR where the quality saturates. In addition, Bjøntegaard model is used to quantify the trade-offs between CR, GoP-size and CSNR, depending on the intended application. Finally, the performance of objective metrics compared to the obtained MOS is evaluated. Results show that Multi-Scale SSIM (MS-SSIM), SSIM and Video Multimethod Assessment Fusion (VMAF) metrics offer the best correlation with the MOS values. Anthony Trioux, Giuseppe Valenzise, Marco Cagnazzo, Michel Kieffer, François-Xavier Coudoux, Patrick Corlay, Mohamed Gharbi |
VCIP | 3 |
| 2020 | Channel Impulsive Noise Mitigation for Linear Video Coding SchemesabstractThis paper considers the problem of impulse noise mitigation when video is encoded using a SoftCast-based Linear Video Coding (LVC) scheme and transmitted using an Orthogonal Frequency-Division Multiplexing (OFDM) scheme for multi-carrier modulation over a wideband channel prone to impulse noise. In the time domain, the impulse noise is modeled as realization of a sequence of independent and identically distributed Bernoulli-Gaussian variables. A Fast Bayesian Matching Pursuit algorithm is employed for impulse noise mitigation. This approach requires the provisioning of some OFDM subchannels to estimate the impulse noise locations and amplitudes. Provisioned subchannels cannot be used to transmit data and lead to a decrease of the video quality at receivers in absence of impulse noise. Using a phenomenological model (PM) of the residual noise variance after impulse mitigation in the subchannels, we have proposed an algorithms that is able to get the amount of subchannel to provision which minimizes the mean-square error of the decoded video at receivers. Simulation results show that the PM can accurately predict the number of subchannels to provision and that impulse noise mitigation can significantly improve the decoded video quality compared to a situation where all subchannels are used for data transmission. Marco Cagnazzo, Michel Kieffer |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Compression Improvement via Reference Organization for 2D-multiview ContentabstractOne of the most challenging goals of future immersive services is to enable the observation of a scene from any viewpoint, thus making free-navigation possible under certain constraints. In order to provide such kind of services with smooth navigation, a huge amount of views should be available on the client's device. In particular, it is important for the case of 2D-multiview content, where cameras are positioned on a 2D grid in order to provide both horizontal and vertical parallax. This kind of content requires a large coding rate; therefore improving the compression performance of video encoders is especially relevant in this case. This paper studies how the encoder configuration affects the compression, by taking into account the spatial position of each camera. Four parameters are addressed in this work: coding order of the views, the number of reference lists, the number of reference pictures, and the ordering of pictures in the reference lists. An average of 12.0% bitrate saving is achieved for medium bitrate and 11.1% for low bitrate compared to the state of the art techniques. Pavel Nikitin, Marco Cagnazzo, Joël Jung |
ICASSP | 2 |
| 2019 | Enhancing HEVC Spatial Prediction by Context-based LearningabstractDeep generative models have been recently employed to compress images, image residuals or to predict image regions. Based on the observation that state-of-the-art spatial prediction is highly optimized from a rate-distortion point of view, in this work we study how learning-based approaches might be used to further enhance this prediction. To this end, we propose an encoder-decoder convolutional network able to reduce the energy of the residuals of HEVC intra prediction, by leveraging the available context of previously decoded neigh-boring blocks. The proposed context-based prediction enhancement (CBPE) scheme enables to reduce the mean square error of HEVC prediction by 25% on average, without any additional signalling cost in the bitstream. Attilio Fiandrotti, Andrei I. Purica, Giuseppe Valenzise, Marco Cagnazzo |
ICASSP | 5 |
| 2019 | Channel Impulsive Noise Mitigation for Linear Video Coding SchemesabstractThis paper considers the problem of impulse noise mitigation for videos encoded using a SoftCast-based Linear Video Coding (LVC) scheme and transmitted using an OFDM scheme over a wideband channel prone to impulse noise. In the time domain, the impulse noise is modeled as realizations of iid Bemoulli-Gaussian variables. A Fast Bayesian Matching Pursuit algorithm is employed for impulse noise mitigation. This approach requires the provisioning of some OFDM subchannels to estimate the impulse noise locations and amplitudes. Provisioned subchannels cannot be used to transmit data and lead to a decrease of the nominal decoded video quality at receivers in absence of impulse noise. Using a phenomenological model (PM) of the residual noise variance after impulse correction, an algorithm is proposed to evaluate the optimal number of subchannels to provision for impulse noise mitigation. Simulation results show that the PM can accurately predict the number of subchannels to provision and that impulse noise mitigation can significantly improve the decoded video quality compared to a situation where all subchannels are used for data transmission. Marco Cagnazzo, Michel Kieffer |
ICASSP | 2 |
| 2019 | Scalable Coding Framework for a View-Dependent Streaming of Digital HologramsabstractUnlike conventional images and videos, digital holograms contain large amounts of data with very low redundancy. Consequently, current communication networks may be not able to meet the bandwidth requirements for hologram transmission in reasonable time. To enable practical streaming of holographic contents, we propose a progressive coding method that combines quality scalability with viewpoint scalability. From a Gabor wavelets decomposition of the hologram, the server starts by selecting the coefficients corresponding to the user's viewpoint. Then, the selected coefficients are encoded progressively according to their importance for the reconstructed view. Experimental results reveal that our approach outperforms conventional scalable codecs and enables the streaming of holographic data with a better quality of experience. Anas El Rhammad, Patrick Gioia, Antonin Gilles, Marco Cagnazzo |
ICIP | 4 |
| 2019 | Optimal and suboptimal channel precoding and decoding matrices for linear video coding
Marco Cagnazzo, Michel Kieffer |
Signal Process. Image Commun. | 2 |
| 2019 | Very Low Bitrate Semantic Compression of Airplane Cockpit Screen ContentabstractThis paper addresses the problem of encoding the video generated by the screen of an airplane cockpit. As other computer screens, cockpit screens consist of computer-generated graphics often atop a natural background. Existing screen content coding schemes fail notably in preserving the readability of textual information at the low bitrates required in avionic applications. We propose a screen coding scheme where textual information is encoded according to the relative semantics rather than in the pixel domain. The encoder localizes textual information, and the semantics of each character are extracted with a convolutional neural network and predictively encoded. Text is then removed via inpainting, and the residual background video is compressed with a standard codec and transmitted to the receiver together with the text semantics. At the decoder side, text is synthesized using the decoded semantics and superimposed over the decoded residual video recovering the original frame. Our proposed scheme offers two key advantages over a semantics-unaware scheme that encodes text in the pixel domain. First, the text readability at the decoder is not compromised by compression artifacts, whereas the relative bitrate is negligible. Second, removal of high-frequency transform coefficients associated with the inpainted text drastically reduces the bitrate of the residual video. Experiments with real cockpit video sequences show BD-rate gains up to 82% and 69% over a reference H.265/HEVC encoder and its screen content coding extension. Moreover, our scheme achieves quasi-errorless character recognition already at very low bitrates, whereas even HEVC-SCC needs at least three or four times more bitrate to achieve a comparable error rate. Iulia Mitrica, Eric Mercier, Christophe Ruellan, Attilio Fiandrotti, Marco Cagnazzo, Béatrice Pesquet-Popescu |
IEEE Trans. Multim. | 5 |
| 2018 | Precoding Matrix Design in Linear Video CodingabstractLinear video coding (LVC) is a promising alternative to classical video coding when video has to be transmitted to wireless receivers experiencing different and time-varying channel conditions. This paper addresses the LVC channel precoding and decoding matrix design when the transmission channel consists of several sub-channels, each with its own power constraint. Such constraints may be found, e.g., in multi-antenna, DSL, or powerline transmission systems. In a previous paper, it has been shown that this matrix design problem may be addressed by an adaptation to LVC of a multi-level water-filling solution proposed for MIMO channels. Here, two suboptimal low-complexity multi-level water-filling techniques are proposed, with different trade-offs between complexity and efficiency. Extensive simulations show that the suboptimal solutions perform very close to the optimal one, with a sensibly reduced complexity. Marco Cagnazzo, Michel Kieffer |
ICASSP | 2 |
| 2018 | Quality Assessment of Deep-Learning-Based Image CompressionabstractImage compression standards rely on predictive coding, transform coding, quantization and entropy coding, in order to achieve high compression performance. Very recently, deep generative models have been used to optimize or replace some of these operations, with very promising results. However, so far no systematic and independent study of the coding performance of these algorithms has been carried out. In this paper, for the first time, we conduct a subjective evaluation of two recent deep-learning-based image compression algorithms, comparing them to JPEG 2000 and to the recent BPG image codec based on HEVC Intra. We found that compression approaches based on deep auto-encoders can achieve coding performance higher than JPEG 2000, and sometimes as good as BPG. We also show experimentally that the PSNR metric is to be avoided when evaluating the visual quality of deep-learning-based methods, as their artifacts have different characteristics from those of DCT or wavelet-based codecs. In particular, images compressed at low bitrate appear more natural than JPEG 2000 coded pictures, according to a no-reference naturalness measure. Our study indicates that deep generative models are likely to bring huge innovation into the video coding arena in the coming years. Giuseppe Valenzise, Andrei I. Purica, Vedad Hulusic, Marco Cagnazzo |
MMSP | 4 |
| 2017 | AVC to HEVC transcoder based on quadtree limitation
Elie Gabriel Mora, Marco Cagnazzo, Frédéric Dufaux |
Multim. Tools Appl. | 2 |
| 2017 | Rate Allocation in Predictive Video Coding Using a Convex Optimization FrameworkabstractOptimal rate allocation is among the most challenging tasks to perform in the context of predictive video coding, because of the dependencies between frames induced by motion compensation. In this paper, using a recursive rate-distortion model that explicitly takes into account these dependencies, we approach the frame-level rate allocation as a convex optimization problem. This technique is integrated into the recent HEVC encoder, and tested on several standard sequences. Experiments indicate that the proposed rate allocation ensures a better performance (in the rate-distortion sense) than the standard HEVC rate control, and with a little loss with respect to an optimal exhaustive research, which is largely compensated by a much shorter execution time. Aniello Fiengo, Giovanni Chierchia, Marco Cagnazzo, Béatrice Pesquet-Popescu |
IEEE Trans. Image Process. | 3 |
| 2016 | View synthesis based on temporal prediction via warped motion vector fieldsabstractThe demand for 3D content has increased over the last years as 3D displays are now widespread. View synthesis methods, such as depth-image-based-rendering, provide an efficient tool in 3D content creation or transmission, and are integrated in coding solutions for multiview video content such as 3D-HEVC. In this paper, we propose a view synthesis method that takes advantage of temporal and inter-view correlations in multiview video sequences. We use warped motion vector fields computed in reference views to obtain temporal predictions of a frame in a synthesized view and blend them with depth-image-based-rendering synthesis. Our method is shown to bring gains of 0.42dB in average when tested on several multiview sequences. Andrei I. Purica, Marco Cagnazzo, Béatrice Pesquet-Popescu, Frédéric Dufaux, Bogdan Ionescu |
ICASSP | 2 |
| 2016 | Depth map coding with elastic contours and 3D surface predictionabstractDepth maps are typically made of smooth regions separated by sharp edges. Following this rationale, this paper presents a novel coding scheme where depth data is represented by a set of contours defining the various regions together with a compact representation of the values inside each region. The proposed coding scheme is based on elastic curves, which make possible to compactly represent the contours exploiting also the temporal consistency in different frames. A 3D surface prediction algorithm is then used to obtain an accurate estimation of the depth field from the coded contours and a subsampled version of the data. Finally, an ad-hoc coding strategy for the low resolution data and the prediction residuals is presented. Experimental results prove how the proposed approach is able to obtain a very high coding efficiency outperforming the HEVC coder at medium-low bitrates. Marco Calemme, Pietro Zanuttigh, Simone Milani, Marco Cagnazzo, Béatrice Pesquet-Popescu |
ICIP | 4 |
| 2016 | Convex optimization for frame-level rate allocation in MV-HEVCabstractOptimal rate allocation is among the most challenging tasks to perform in the context of multi-view video coding, because of the dependency between frames induced by motion compensation and depth image-based rendering. In this paper, using a recursive rate-distortion model that explicitly takes into account these dependencies, we approach the frame-level rate allocation as a convex optimization problem. Within this framework, we provide an efficient algorithm for exactly solving the above problem with recent convex optimization tools. Experiments on standard sequences demonstrate the interest of considering the proposed rate allocation method and confirm that our approach ensures a better performance (in ratedistortion sense) than the standard MV-HEVC rate control. Aniello Fiengo, Giovanni Chierchia, Marco Cagnazzo, Béatrice Pesquet-Popescu |
ICIP | 3 |
| 2016 | Softcast with per-carrier power-constrained channelsabstractThis paper considers the Softcast joint source-channel video coding scheme for data transmission over parallel channels with different power constraints and noise characteristics, typical in DSL or PLT channels. To minimize the mean square error at receiver, an optimal precoding matrix design problem has to be solved, which requires the solution of an inverse eigenvalue problem. Such solution is taken from the MIMO channel precoder design literature. Alternative suboptimal precoding matrices are also proposed and analyzed, showing the efficiency of the optimal precoding matrix within Softcast, which provides gains increasing with the encoded video quality. Marc Antonini, Marco Cagnazzo, Lorenzo Guerrieri, Michel Kieffer, Irina Delia Nemoianu, Roger Samy |
ICIP | 3 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Multiview Plus Depth Video Coding With Temporal Prediction View SynthesisabstractMultiview video (MVV) plus depths formats use view synthesis to build intermediate views from existing adjacent views at the receiver side. Traditional view synthesis exploits the disparity information to interpolate an intermediate view by considered inter-view correlations. However, temporal correlation between different frames of the intermediate view can be used to improve the synthesis. We propose a new coding scheme for 3-D High Efficiency Video Coding (HEVC) that allows us to take full advantage of temporal correlations in the intermediate view and improve the existing synthesis from adjacent views. We use optical flow techniques to derive dense motion vector fields (MVF) from the adjacent views and then warp them at the level of the intermediate view. This allows us to construct multiple temporal predictions of the synthesized frame. A second contribution is an adaptive fusion method that judiciously selects between temporal and inter-view prediction to eliminate artifacts associated with each prediction type. The proposed system is compared against the state-of-the-art view synthesis reference software 1-D Fast technique used in 3-D HEVC standardization. Three intermediary views are synthesized. Gains of up to 1.21-dB Bjontegaard Delta peak SNR are shown when evaluated on several standard MVV test sequences. Andrei I. Purica, Elie Gabriel Mora, Béatrice Pesquet-Popescu, Marco Cagnazzo, Bogdan Ionescu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Reference View Selection in DIBR-Based Multiview CodingabstractAugmented reality, interactive navigation in 3D scenes, multiview video, and other emerging multimedia applications require large sets of images, hence larger data volumes and increased resources compared with traditional video services. The significant increase in the number of images in multiview systems leads to new challenging problems in data representation and data transmission to provide high quality of experience on resource-constrained environments. In order to reduce the size of the data, different multiview video compression strategies have been proposed recently. Most of them use the concept of reference or key views that are used to estimate other images when there is high correlation in the data set. In such coding schemes, the two following questions become fundamental: 1) how many reference views have to be chosen for keeping a good reconstruction quality under coding cost constraints? And 2) where to place these key views in the multiview data set? As these questions are largely overlooked in the literature, we study the reference view selection problem and propose an algorithm for the optimal selection of reference views in multiview coding systems. Based on a novel metric that measures the similarity between the views, we formulate an optimization problem for the positioning of the reference views, such that both the distortion of the view reconstruction and the coding rate cost are minimized. We solve this new problem with a shortest path algorithm that determines both the optimal number of reference views and their positions in the image set. We experimentally validate our solution in a practical multiview distributed coding system and in the standardized 3D-HEVC multiview coding scheme. We show that considering the 3D scene geometry in the reference view, positioning problem brings significant rate-distortion improvements and outperforms the traditional coding strategy that simply selects key frames based on the distance between cameras. Thomas Maugey, Giovanni Petrazzuoli, Pascal Frossard, Marco Cagnazzo, Béatrice Pesquet-Popescu |
IEEE Trans. Image Process. | 4 |
| 2015 | Improved view synthesis by motion warping and temporal hole fillingabstractView synthesis received increasing attention over the last years, as it offers a wide range of practical applications like Free Viewpoint Television, 3D video, video gaming, etc. The main issues in view synthesis are the filling of disoccluded areas and the warping of real views. In this paper we propose a new hole filling method, it uses temporal correlations in the real views to extract information on disoccluded areas from different time instants in the synthetic view. We also propose a sub-pixel warping technique that takes into account depth and can be used for both the warping of the real view as well as for motion compensation. Our method is proved to bring gains of up to 0.31dB in average over several multiview test sequences. Andrei I. Purica, Elie Gabriel Mora, Béatrice Pesquet-Popescu, Marco Cagnazzo, Bogdan Ionescu |
ICASSP | 4 |
| 2015 | Shannon-Kotelnikov mappings for softcast-based joint source-channel video codingabstractThis paper introduces Shannon-Kotelnikov (SK) mapping in the SoftCast joint source-channel video coding scheme. On bandwidth constrained channels, the performance of SoftCast saturates, due to the large amount of data (chunks) dropped to match the bandwidth requirements. Using SK mapping, it is possible to increase the number of chunks that may be transmitted without increasing the bandwidth requirements. The resulting scheme has an increased number of design parameters for which we present a transmission-power constrained optimization. This extends range of channel SNRs over which the PSNR gracefully increases and improves the end-to-end performance at medium to high SNRs. The price to be paid is a performance degradation at low SNRs. Marco Cagnazzo, Michel Kieffer |
ICIP | 1 |
| 2015 | ROI-based rate control using tiles for an HEVC encoded video stream over a lossy networkabstractThe growth in the use of high definition (HD) and above video resolutions streams has outstripped the rate at which network infrastructure has been deployed. Video streaming applications require appropriate rate control techniques that make use of the specific characteristics of the video content, such as the regions of interest (ROI). With the introduction of high efficiency video coding (HEVC) streams, we consider new coding features to make a novel ROI-based rate control (RC) algorithm. The proposed approach introduces tiling in a ROI-based rate control scheme. It aims at enhancing the quality of important regions (i.e. faces for a videoconferencing system) considering independently coded regions lying within an ROI and helps evaluating the ROI quality under poor channel conditions. Our work consists of two major steps. First, we designed a RC algorithm based on an independent processing of tiles of different regions. Second, we investigate the effect of ROI- and tile-based rate control algorithm on the decoded quality of the stream transmitted over a lossy channel. Marwa Meddeb, Marco Cagnazzo, Béatrice Pesquet-Popescu |
ICIP | 2 |
| 2015 | Contour-Based Depth Coding: A Subjective Quality Assessment StudyabstractMulti-view video plus depth is emerging as the most flexible format for 3D video representation, as witnessed by the current standardization efforts by ISO and ITU. The depth information allows synthesizing virtual view points, and for its compression various techniques have been proposed. It is generally recognized that a high quality view rendering at the receiver side is possible only by preserving the contour information since distortions on edges during the encoding step would cause a sensible degradation on the synthesized view and on the 3D perception. As a consequence recent approaches include contour-based coding of depths. However, the impact of contour-preserving depth-coding on the perceived quality of synthesized images has not been conveniently studied. Therefore in this paper we make an investigation by means of a subjective study to better understand the limits and the potentialities of the different techniques. Our results show that the contour information is indeed relevant in the synthesis step: preserving the contours and coding coarsely the rest typically leads to images that users cannot tell apart from the reference ones, even at low bit rate. Moreover, our results show that objective metrics that are commonly used to evaluate synthesized images may have a low correlation coefficient with MOS rates and are in general not consistent across several techniques and contents. Marco Calemme, Marco Cagnazzo, Béatrice Pesquet-Popescu |
ISM | 2 |
| 2015 | Subjectie evaluation of Super Multi-View compressed contents on high-end light-field 3D displays
Antoine Dricot, Joël Jung, Marco Cagnazzo, Béatrice Pesquet-Popescu, Frédéric Dufaux, Péter Tamás Kovács, Vamsi Kiran Adhikarla |
Signal Process. Image Commun. | 3 |
| 2015 | Fusion of Global and Local Motion Estimation Using Foreground Objects for Distributed Video CodingabstractThe side information (SI) in Distributed Video Coding (DVC) is estimated using the available decoded frames and exploited for the decoding and reconstruction of other frames. The quality of the SI has a strong impact on the performance of DVC. Here, we propose a new approach that combines both global and local SI to improve coding performance. Since the background pixels in a frame are assigned to global estimation and the foreground objects to local estimation, one needs to estimate foreground objects in the SI using the backward and forward foreground objects, the background pixels are directly taken from the global SI. Specifically, elastic curves and local motion compensation are used to generate the foreground objects masks in the SI. Experimental results show that, as far as the rate-distortion performance is concerned, the proposed approach can achieve a PSNR improvement of up to 1.39 dB for a group of picture (GOP) size of 2, and up to 4.73 dB for larger GOP sizes, with respect to the reference DISCOVER codec. Abdalbassir Abou-Elailah, Frédéric Dufaux, Joumana Farah, Marco Cagnazzo, Anuj Srivastava, Béatrice Pesquet-Popescu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | A convex-optimization framework for frame-level optimal rate allocation in predictive video codingabstractOptimal rate allocation is among the most challenging tasks to perform in the context of predictive video coding, because of the dependencies between frames induced by motion compensation. In this paper, we derive an analytical rate-distortion model that explicitly takes into account the dependencies between frames. The proposed approach allows us to formulate the frame-level optimal rate allocation as a convex optimization problem. Within this framework, we are able to achieve the exact solution in limited time (even for large-size problems), thanks to the flexibility offered by recent convex optimization techniques. Experiments on standard sequences demonstrate the interest of considering the proposed rate-distortion model and confirm that the optimal rate allocation ensures a better distribution of the total bit budget, with superior results (in the rate-distortion sense) with respect to the standard H.264/AVC rate control. Aniello Fiengo, Giovanni Chierchia, Marco Cagnazzo, Béatrice Pesquet-Popescu |
ICASSP | 3 |
| 2014 | Region-of-interest based rate control scheme for high efficiency video codingabstractIn this paper, we propose a new rate control scheme designed for the newest high efficiency video coding (HEVC) standard, and aimed at enhancing the quality of regions of interest (ROI). Our approach allocates a higher bit rate to the region of interest while keeping the global bit rate close to the assigned target value. This algorithm is developed for a videoconferencing system, where the ROIs (typically, faces) are automatically detected and each coding unit is classified in a region of the interest map. This map is given as input to the rate control algorithm and the bit allocation is made accordingly. Experimental results show that the proposed scheme achieves accurate target bit rates and provides an improvement in the region of interest quality, both in objective metrics and based on subjective quality evaluation. Marwa Meddeb, Marco Cagnazzo, Béatrice Pesquet-Popescu |
ICASSP | 2 |
| 2014 | Full parallax super multi-view video codingabstractSuper Multi-View (SMV) video is a key enabler for future 3D video services that allows a glasses-free visualization and eliminates many causes of discomfort existing in current available 3D video technologies. SMV video content is composed of tens or hundreds of views, that can be aligned in horizontal only or both horizontal and vertical directions, providing respectively horizontal parallax or full parallax. This paper compares several coding schemes and coding orders, and proposes a coding structure that exploits inter-view correlations in the two directions, providing BD-rate gains up to 29.1% when compared to a basic anchor structure. Additionally, Neighboring Block Disparity Vector (NBDV) and Inter-View Motion Prediction (IVMP) coding tools are further improved to efficiently exploit coding structures in two dimensions, with BD-rate gains up to 4.2% reported over the reference 3D-HEVC encoder. Antoine Dricot, Joël Jung, Marco Cagnazzo, Béatrice Pesquet-Popescu, Frédéric Dufaux |
ICIP | 3 |
| 2014 | Key view selection in distributed multiview codingabstractMultiview image and video systems with large number of views lead to new problems in data representation, transmission and user interaction. In order to reduce the data volumes, most distributed multiview coding schemes exploit the inter-view redundancies at the decoder side, using view synthesis from key views. In the situation where many views are considered, the two following questions become fundamental: i) how many key views have to be chosen for keeping a good reconstruction quality with reasonable coding cost? ii) where to place them optimally in the multiview sequences? We propose in this paper an algorithm for selecting the key views in a distributed multiview coding scheme. Based on a novel metric for the correlation between the views, we formulate an optimization problem for the positioning of the key views such that both the distortion of the reconstruction and the coding rate cost are effectively minimized. We then propose a new optimization strategy based on shortest path algorithm that permits to determine both the optimal number of key views and their positions in the image set. We experimentally validate our solution in a practical distributed multiview coding system and we show that considering the 3D scene geometry in the key view positioning brings significant rate-distortion improvements compared to distance-based key view selection as it is commonly done in the literature. Thomas Maugey, Giovanni Petrazzuoli, Pascal Frossard, Marco Cagnazzo, Béatrice Pesquet-Popescu |
VCIP | 4 |
| 2014 | Initialization, Limitation, and Predictive Coding of the Depth and Texture Quadtree in 3D-HEVCabstractThe 3D video extension of High Efficiency Video Coding (3D-HEVC) exploits texture-depth redundancies in 3D videos using intercomponent coding tools. It also inherits the same quadtree coding structure as HEVC for both components. The current software implementation of 3D-HEVC includes encoder shortcuts that speed up the quadtree construction process, but those are always accompanied by coding losses. Furthermore, since the texture and its associated depth represent the same scene, at the same time instant and view point, their quadtrees are closely linked. In this paper, an intercomponent tool is proposed in which this link is exploited to save both runtime and bits through a joint coding of the quadtrees. If depth is coded before the texture, the texture quadtree is initialized from the coded depth quadtree. Otherwise, the depth quadtree is limited to the coded texture quadtree. A 31% encoder runtime saving, a -0.3% gain for coded and synthesized views and a -1.8% gain for coded views are reported for the second method. Elie Gabriel Mora, Joël Jung, Marco Cagnazzo, Béatrice Pesquet-Popescu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | On a Hashing-Based Enhancement of Source Separation Algorithms Over Finite Fields With Network Coding PerspectivesabstractBlind Source Separation (BSS) deals with the recovery of source signals from a set of observed mixtures, when little or no knowledge of the mixing process is available. BSS can find an application in the context of network coding, where relaying linear combinations of packets maximizes the throughput and increases the loss immunity. By relieving the nodes from the need to send the combination coefficients, the overhead cost is largely reduced. However, the scaling ambiguity of the technique and the quasi-uniformity of compressed media sources makes it unfit, at its present state, for multimedia transmission. In order to open new practical applications for BSS in the context of multimedia transmission, we have recently proposed to use a non-linear encoding to increase the discriminating power of the classical entropy-based separation methods. Here, we propose to append to each source a non-linear message digest, which offers an overhead smaller than a per-symbol encoding and that can be more easily tuned. Our results prove that our algorithm is able to provide high decoding rates for different media types such as image, audio, and video, when the transmitted messages are less than 1.5 kilobytes, which is typically the case in a realistic transmission scenario. Irina Delia Nemoianu, Claudio Greco 0001, Marco Cagnazzo, Béatrice Pesquet-Popescu |
IEEE Trans. Multim. | 3 |
| 2014 | Depth-Based Multiview Distributed Video CodingabstractMultiview distributed video coding (DVC) has gained much attention in the last few years because of its potential in avoiding communication between cameras without decreasing the coding performance. However, the current results are not matching the expectations mainly due to the fact that some theoretical assumptions are not satisfied in the current implementations. For example, in distributed source coding the encoder must know the correlation between the sources, which cannot be achieved in the traditional DVC systems without having a communication between the cameras. In this work, we propose a novel multiview distributed video coding scheme in which the depth maps are used to estimate the way two views are correlated with no exchanges between the cameras. Only their relative positions are known. We design the complete scheme and further propose a rate allocation algorithm to efficiently share the bit budget between the different components of our scheme. Then, a rate allocation algorithm for depth maps is proposed in order to maximize the quality of synthesized virtual views. We show, through detailed experiments, that our scheme significantly outperforms the state-of-the-art DVC system. Giovanni Petrazzuoli, Thomas Maugey, Marco Cagnazzo, Béatrice Pesquet-Popescu |
IEEE Trans. Multim. | 3 |
| 2013 | On a practical approach to source separation over finite fields for network coding applicationsabstractIn Blind Source Separation, or BSS, a set of source signals are recovered from a set of mixed observations without knowledge of the mixing parameters. Originated for real signals, BSS has recently been applied to finite fields, enabling more practical applications. However, classical entropy-based techniques do not perform well in finite fields. Here, we propose a non-linear encoding of the sources to increase the discriminating power of the separation methods. Our results show that the encoding improves the success rate of the separation for sources with few samples in large finite fields, both conditions met in practical networking applications. Our results open new possibilities in the context of network coding-wherein linear combinations of packets are sent in order to maximize throughput and increase loss immunity- by relieving the nodes from the need to send the combination coefficients, thus reducing the overhead cost. Irina Delia Nemoianu, Claudio Greco 0001, Marc Castella, Béatrice Pesquet-Popescu, Marco Cagnazzo |
ICASSP | 5 |
| 2013 | Modification of the merge candidate list for dependent views in 3D-HEVCabstractA test model for an HEVC-based 3D video coding standard (3D-HEVC) has recently been drafted. 3D-HEVC exploits inter-view redundancies by including disparity-compensated prediction (DCP) for efficient dependent view coding. It also uses the Merge coding mode to reduce the cost of motion / disparity parameters. However, the candidates in the Merge list are mostly temporal motion vectors. DCP does not often benefit from accurate predictors and is thus costly. Consequently, motion-compensated prediction (MCP) remains largely preferred. In this paper, we propose to reduce the cost of DCP by modifying the Merge candidate list to always include a disparity vector candidate. Two methods are proposed: the new candidate is either added in the secondary or in the primary list of candidates. The latter method, which achieves average bitrate reductions of 0.6% for dependent views, and 0.2% for coded and synthesized views, was adopted in both the 3D-HEVC working draft and software. Elie Gabriel Mora, Joël Jung, Marco Cagnazzo, Béatrice Pesquet-Popescu |
ICIP | 3 |
| 2013 | Modification of the disparity vector derivation process in 3D-HEVCabstractThe up-and-coming extension of HEVC for 3D video (3D-HEVC) includes various tools to exploit different redundancies in a 3D video signal. Inter-view redundancies are in particular exploited using Inter-View Motion Prediction (IVMP) and Inter-View Residual Prediction (IVRP). Both of these tools compensate disparity-wise the current prediction unit (PU) in order to find its corresponding PU in a base view, from which some prediction information for the current PU is retrieved. The disparity vector (DV) used for disparity compensation is currently derived using a neighboring search process (NBDV) for a DV across spatial and temporal neighbors. The first DV found is selected as the final DV used in IVMP and IVRP, with no guarantee of optimality. In this paper, the NBDV derivation process is changed: all found DVs from different neighbors are stored in a list. Redundant vectors in this list are removed, and a median computation on the remaining vectors is performed. The resulting DV is set as the DV used for IVMP. Average bitrate reductions of 0.6% and 0.8% for the two dependent views and 0.2% on synthesized views are reported with only a slight increase in encoder and decoder runtimes. Elie Gabriel Mora, Joël Jung, Béatrice Pesquet-Popescu, Marco Cagnazzo |
MMSP | 4 |
| 2013 | Fusion of Global and Local Motion Estimation for Distributed Video CodingabstractThe quality of side information plays a key role in distributed video coding. In this paper, we propose a new approach that consists of combining global and local motion compensation at the decoder side. The parameters of the global motion are estimated at the encoder using scale invariant feature transform features. Those estimated parameters are sent to the decoder in order to generate a globally motion compensated side information. Conversely, a locally motion compensated side information is generated at the decoder based on motion-compensated temporal interpolation of neighboring reference frames. Moreover, an improved fusion of global and local side information during the decoding process is achieved using the partially decoded Wyner-Ziv frame and decoded reference frames. The proposed technique improves significantly the quality of the side information, especially for sequences containing high global motion. Experimental results show that, as far as the rate-distortion performance is concerned, the proposed approach can achieve a PSNR improvement of up to 1.9 dB for a Group of Pictures (GOP) size of 2, and up to 4.65 dB for larger GOP sizes, with respect to the reference DISCOVER codec. Abdalbassir Abou-Elailah, Frédéric Dufaux, Joumana Farah, Marco Cagnazzo, Béatrice Pesquet-Popescu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2013 | Evaluation of Side Information Effectiveness in Distributed Video CodingabstractThe rate-distortion performance of a distributed video coding system strongly depends on the characteristics of the side information. One could naïvely think that the best side information is the one with the largest PSNR with respect to the original corresponding image. However, previous works have shown that this is not always the case and a reduction of the side information MSE does not always translate into better rate-distortion performance for the complete system. The scope of this paper is to explore a set of metrics other than the PSNR and explicitly designed to classify the side information with respect to its impact on the end-to-end compression performance. A first contribution is to define an experimental framework that can be used to meaningfully compare different metrics for side information evaluation. As a second contribution, our analysis allows to understand why in some cases PSNR-based metrics provide a fairly reliable estimation of the side information quality, while in other cases they do not. This analysis also allows us to introduce a set of new metrics that are better adapted for side information effectiveness evaluation, and that are based on a suitable power of the absolute difference between side information and the original image, or on the Hamming distance between the respective transform coefficients. Besides their theoretical interest, these new metrics can also improve the rate-distortion performance of some distributed video coding systems such as the hash-based ones. We observe improvement up to 74% rate reduction in a simple study case. Thomas Maugey, Jérôme Gauthier, Marco Cagnazzo, Béatrice Pesquet-Popescu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | A framework for joint multiple description coding and network coding over wireless ad-hoc networksabstractNetwork coding (NC) can achieve the maximum information flow in the network by allowing nodes to combine received packets before retransmission. Several papers have shown NC to be beneficial in mobile ad-hoc networks, but the delay introduced by buffered decoding raises a problem in real-time streaming applications. Here we propose to use NC jointly with multiple description coding (MDC) to allow instant decoding of the received packets. The optimal encoding coefficients are chosen via distributed optimisation of the expected video quality. Nodes receive up-to-date information about the network topology through a recently proposed protocol, originally designed for real-time streaming of MDC video. Results show that, due to the limitations imposed by instant decoding to the coding window size, our approach consistently outperforms the popular technique of random linear network coding. Irina Delia Nemoianu, Claudio Greco 0001, Marco Cagnazzo, Béatrice Pesquet-Popescu |
ICASSP | 3 |
| 2012 | Motion prediction of depth video for depth-image-based rendering using don't care regionsabstractTo enable synthesis of any desired intermediate view between two captured views at decoder via depth-image-based rendering (DIBR), both texture and depth maps from the captured viewpoints must be encoded and transmitted in a format known as texture-plus-depth. In this paper, we focus on the compression of depth maps across time to lower the overall bitrate in texture-plus-depth format. We observe that depth maps are not directly viewed, but are only used to provide geometric information of the captured scene for view synthesis at decoder. Thus, as long as the resulting geometric error does not lead to unacceptable synthesized view quality, each depth pixel only needs to be reconstructed at the decoder coarsely within a tolerable range. We first formalize the notion of tolerable range per depth pixel as don't care region (DCR), by studying the synthesized view distortion sensitivity to the pixel value - a sensitive depth pixel will have a narrow DCR, and vice versa. Given per-pixel DCRs, we then modify inter-prediction modes during motion prediction to search for a predictor block matching per-pixel DCRs in a target block (rather than the fixed ground truth depth signal in a target block), in order to lower the energy of the prediction residual for the block. We implemented our DCR-based motion prediction scheme inside H.264; our encoded bitstreams remain 100% standard compliant. We show experimentally that our proposed encoding scheme can reduce the bitrate of depth maps coded with baseline H.264 by over 28%. Giuseppe Valenzise, Gene Cheung, Rafael Galvão de Oliveira, Marco Cagnazzo, Béatrice Pesquet-Popescu, Antonio Ortega |
PCS | 4 |
| 2012 | Multi-view video streaming over wireless networks with RD-optimized scheduling of network coded packetsabstractMulti-view video streaming is an emerging video paradigm that enables new interactive services, such as free viewpoint television and immersive teleconferencing. However, it comes with a high bandwidth cost, as the equivalent of many single-view streams has to be transmitted. Network coding (NC) can improve the performance of the network by allowing nodes to combine received packets before retransmission. Several works have shown NC to be beneficiai in wireless networks, but the delay introduced by buffering before decoding raises a problem in real-time streaming applications. Here, we propose to use Expanding Window NC (EWNC) for multi-view streaming to allow immediate decoding of the received packets. The order in which the packets are included in the coding window is chosen via RD-optimization for the current sending opportunity. Results show that our approach consistently outperforms both classical NC applied on each view independently and transmission without NC. Irina Delia Nemoianu, Claudio Greco 0001, Marco Cagnazzo, Béatrice Pesquet-Popescu |
VCIP | 3 |
| 2012 | Low-Latency Video Streaming With Congestion Control in Mobile Ad-Hoc NetworksabstractIn this paper, we address the challenge of delivering a video stream, encoded with multiple descriptions, in a mobile ad-hoc environment with low-latency constraints. This kind of application is meant to provide an efficient and reliable video communication tool in scenarios where the deployment of an infrastructure is not feasible, such as military and disaster relief applications. First, we present a recently proposed protocol that employs a reliable form of one-hop broadcast to build an efficient overlay network according to a multi-objective function that minimizes the number of packets injected in the network and maximizes the path diversity among descriptions. Then, we introduce the main contribution of this paper: a cross-layer congestion control strategy where the MAC layer is video-coding aware and adjusts its transmission parameters (namely, the RTS retry limit) via congestion/distortion optimization. The main challenge in this approach is providing a reliable estimation of congestion and distortion, given the limited information available at each node. Our simulations show that, if a stringent constraint of low delay is imposed, our technique grants a consistent gain in terms of both PSNR and delay reduction, for bitrates up to a few megabits per second. Claudio Greco 0001, Marco Cagnazzo, Béatrice Pesquet-Popescu |
IEEE Trans. Multim. | 2 |
| 2011 | Using distributed source coding and depth image based rendering to improve interactive multiview video accessabstractMultiple-views video is commonly believed to be the next significant achievement in video communications, since it enables new exciting interactive services such as free viewpoint television and immersive teleconferencing. However the interactivity requirement (i.e. allowing the user to change the viewpoint during video streaming) involves a trade-off between storage and bandwidth costs. Several solutions have been proposed in the literature, using redundant predictive frames, Wyner-Ziv frames, or a combination of them. In this paper, we adopt distributed video coding for interactive multiview video plus depth (MVD), taking advantage of depth image based rendering (DIBR) and depth-aided inpainting to fill the occlusion areas. To the authors' best knowledge, very few works in interactive MVD consider the problem of continuity of the playback during the switching among streams. Therefore we survey the existing solutions, we propose a set of techniques for MVD coding and we compare them. As main results, we observe that DIBR can help in rate reduction (up to 13.36% for the texture video and up to 8.67% for the depth map, wrt the case where DIBR is not used), and we also note that the optimal strategy to combine DIBR and distributed video coding depends on the position of the switching time into the group of pictures. Choosing the best technique on a frame-to-frame basis can further reduce the rate from 1% to 6%. Giovanni Petrazzuoli, Marco Cagnazzo, Frédéric Dufaux, Béatrice Pesquet-Popescu |
ICIP | 2 |
| 2011 | Wyner-ziv coding for depth maps in multiview video-plus-depthabstractThree dimensional digital video services are gathering a lot of attention in recent years, thanks to the introduction of new and efficient acquisition and rendering devices. In particular, 3D video is often represented by a single view and a so called depth map, which gives information about the distance between the point of view and the objects. This representation can be extended to multiple views, each with its own depth map. Efficient compression of this kind of data is of course a very important topic in sight of a massive deployment of services such as 3D-TV and FTV (free viewpoint TV). In this paper we consider the application of distributed coding techniques to the coding of depth maps, in order to reduce the complexity of single view or multi view encoders and to enhance interactive multiview video streaming. We start from state-of-the-art distributed video coding techniques and we improve them by using high order motion interpolation and by exploiting texture motion information to encode the depth maps. The experiments reported here show that the proposed method achieves a rate reduction up to 11.06% compared to state-of-the-art distributed video coding technique. Giovanni Petrazzuoli, Marco Cagnazzo, Frédéric Dufaux, Béatrice Pesquet-Popescu |
ICIP | 2 |
| 2011 | An MDC-based video streaming architecture for mobile networksabstractMultiple description coding (MDC) is a framework designed to improve the robustness of video content transmission in lossy environments. In this work, we propose an MDC technique using a legacy coder to produce two descriptions, based on separation of even and odd frames. If only one description is received, the missing frames are reconstructed using temporal high-order motion interpolation (HOMI), a technique originally proposed for distributed video coding. If both descriptions are received, the frames are reconstructed as a block-wise linear combination of the two descriptions, with the coefficient computed at the encoder in a RD-optimised fashion, encoded with a context-adaptive arithmetic coder, and sent as side information. We integrated the proposed technique in a mobile ad-hoc streaming protocol, and tested it using a group mobility model. The results show a non-negligible gain for the expected video quality, with respect to the reference technique. Claudio Greco 0001, Giovanni Petrazzuoli, Marco Cagnazzo, Béatrice Pesquet-Popescu |
MMSP | 3 |
| 2011 | A New Coding Mode for Hybrid Video Coders Based on Quantized Motion VectorsabstractThe rate allocation tradeoff between motion vectors and transform coefficients has a major importance when it comes to efficient video compression. This paper introduces a new coding mode for an H.264/AVC-like video coder, which improves the management of this resource allocation. The proposed technique can be used within any hybrid video encoder allowing a different coding mode for any macroblock. The key tool of the new mode is the lossy coding of motion vectors, obtained via quantization: while the transformed motion-compensated residual is computed with a high-precision motion vector, the motion vector itself is quantized before being sent to the decoder, in a rate/distortion optimized way. Several problems have to be faced with in order to get an efficient implementation of the coding mode, especially the coding and prediction of the quantized motion vectors, and the selection and encoding of the quantization steps. This new coding mode improves the performance of the hybrid video encoder over several sequences at different resolutions. Marie Andrée Agostini, Marco Cagnazzo, Marc Antonini, Guillaume Laroche, Joël Jung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | High order motion interpolation for side information improvement in DVCabstractA key step in distributed video coding is the generation of the side information (SI) i.e. the estimation of the Wyner-Ziv frame (WZF). This step is also frequently called image interpolation. State-of-the-art techniques perform a motion estimation between adjacent key frames (KFs) and linear interpolation in order to assess object positions in the WZF, and then the SI is produced by motion compensating the KFs. However the uniform motion model underlying this approach is not always able to produce a satisfying estimation of the motion, which can result in a low SI quality. In this paper we propose a new method for the generation of SI, based on higher order motion interpolation. We use more than two KFs to estimate the position of the current WZF block, which allows us to correctly estimate more complex motion (such as, for example, uniform accelerated motion). We performed a number of tests for the fine tuning of the parameters of the method. Our experiments show that the new interpolation technique has a small computational cost increase with respect to state of the art, but provides remarkably better performance with up to 0.5 dB of PSNR improvement in SI quality. Moreover the proposed method performs consistently well for several GOP sizes. Giovanni Petrazzuoli, Marco Cagnazzo, Béatrice Pesquet-Popescu |
ICASSP | 2 |
| 2010 | Robust decoding of a 3D-ESCOT bitstream transmitted over a noisy channelabstractIn this paper, we propose a joint source-channel (JSC) decoding scheme for 3D ESCOT-based video coders, such as Vidwav. The embedded bitstream generated by such coders is very sensitive to transmission errors unavoidable on wireless channels. The proposed JSC decoder employs the residual redundancy left in the bitstream by the source coder combined with bit reliability information provided by the channel or channel decoder to correct transmission errors. When considering an AWGN channel, the performance gains are in average 4 dB in terms of PSNR of the reconstructed frames, and 0.7 dB in terms of channel SNR. When considering individual frames, the obtained gain is up to 15 dB in PSNR. Manel Abid, Michel Kieffer, Marco Cagnazzo, Béatrice Pesquet-Popescu |
ICIP | 3 |
| 2010 | H.264-based multiple description coding using motion compensated temporal interpolationabstractMultiple description coding is a framework adapted to noisy transmission environments. In this work, we use H.264 to create two descriptions of a video sequence, each of them assuring a minimum quality level. If both of them are received, a suitable algorithm is used to produce an improved quality sequence. The key technique is a temporal image interpolation using motion compensation, inspired to the distributed video coding context. The interpolated image blocks are weighted with the received blocks obtained from the other description. The optimal weights are computed at the encoder and efficiently sent to the decoder as side information. The proposed technique shows a remarkable gain for central decoding with respect to similar methods available in the state of the art. Claudio Greco 0001, Marco Cagnazzo, Béatrice Pesquet-Popescu |
MMSP | 2 |
| 2010 | Side information enhancement using an adaptive hash-based genetic algorithm in a Wyner-Ziv contextabstractSide information construction in Wyner-Ziv video coding is a sensible task which strongly influences the final ratedistortion performance of the scheme. This side information is usually generated through an interpolation of the previous and next images. Some of the zones of a scene however, such as the occlusions, cannot be estimated with other frames. In this paper we propose to avoid this problem by sending some hash information for these unpredictable zones of the image. The resulting algorithm is described and tested here. The obtained results show the advantages of using localized hash information for the high error zones in distributed video coding. Thomas Maugey, Charles Yaacoub, Joumana Farah, Marco Cagnazzo, Béatrice Pesquet-Popescu |
MMSP | 4 |
| 2010 | Side information refinement for long duration GOPs in DVCabstractSide information generation is a critical step in distributed video coding systems. This is performed by using motion compensated temporal interpolation between two or more key frames (KFs). However, when the temporal distance between key frames increases (i.e. when the GOP size becomes large), the linear interpolation becomes less effective. In a previous work we showed that this problem can be mitigated by using high order interpolation. Now, in the case of long duration GOP, state-of-the-art algorithms propose a hierarchical algorithm for side information generation. By using this procedure, the quality of the central interpolated image in a GOP is consistently worse than images closer to the KFs. In this paper we propose a refinement of the central WZFs by higher order interpolation of the already decoded WZFs, that are closer to the WZF to be estimated. So we reduce the fluctuation of side information quality, with a beneficial impact on final rate-distortion characteristics of the system. The experimental results show an improvement on the SI up to 2.71 dB with respect the state-of-the-art and a global improvement of the PSNR on the decoded frames up to 0.71 dB and a bit rate reduction up to 15%. Giovanni Petrazzuoli, Thomas Maugey, Marco Cagnazzo, Béatrice Pesquet-Popescu |
MMSP | 3 |
| 2010 | Introducing differential motion estimation into hybrid video codersabstractDifferential motion estimation produces dense motion vector fields which are far too demanding in terms of coding rate in order to be used in video coding. However, a pel-recursive technique like that introduced by Cafforio and Rocca can be modified in order to work using only the information available at the decoder side. This allows to improve the motion vectors produced in the classical predictive modes of H.264. In this paper we describe the modification needed in order to introduce a differential motion estimation method into the H.264 codec. Experimental results will validate a coding mode, opening new perspectives in using differential-based motion estimation techniques into classical hybrid codecs. Marco Cagnazzo, Béatrice Pesquet-Popescu |
VCIP | 1 |
| 2010 | Mutual information-based context quantization
Marco Cagnazzo, Marc Antonini, Michel Barlaud |
Signal Process. Image Commun. | 1 |
| 2009 | A differential motion estimation method for image interpolation in distributed video codingabstractMotion estimation methods based on differential techniques proved to be very useful in the context of video analysis, but have a limited employment in classical video compression because, though accurate, the dense motion vector field they produce requires too much coding resource and computational effort. On the contrary, this kind of algorithm could be useful in the framework of distributed video coding (DVC). In this paper we propose a differential motion estimation algorithm which can run at the decoder in a DVC scheme, without requiring any increase in coding rate. This algorithm allows a performance improvement in image interpolation with respect to state-of-the-art algorithms. Marco Cagnazzo, Thomas Maugey, Béatrice Pesquet-Popescu |
ICASSP | 1 |
| 2009 | Image interpolation with edge-preserving differential motion refinementabstractMotion estimation (ME) methods based on differential techniques provide useful information for video analysis, and moreover it is relatively easy to embed into them regularity constraints enforcing for example, contour preservation. On the other hand, these techniques are rarely employed for video compression since, though accurate, the dense motion vector field (MVF) they produce requires too much coding resource and computational effort. However, this kind of algorithm could be useful in the framework of distributed video coding (DVC), where the motion vector are computed at the decoder side, so that no bit-rate is needed to transmit them. Moreover usually the decoder has enough computational power to face with the increased complexity of differential ME. In this paper we introduce a new image interpolation algorithm to be used in the context of DVC. This algorithm combines a popular DVC technique with differential ME. We adapt a pel-recursive differential ME algorithm to the DVC context; moreover we insert a regularity constraint which allows more consistent MVFs. The experimental results are encouraging: the quality of interpolated images is improved of up to 1.1 dB w.r.t. to state-of-the-art techniques. These results prove to be consistent when we use different GOP sizes. Marco Cagnazzo, Wided Miled, Thomas Maugey, Béatrice Pesquet-Popescu |
ICIP | 1 |
| 2009 | Adaptive video streaming with long term feedbacksabstractThis paper proposes a video streaming system optimizing resource utilization when the media server only disposes of long term feedbacks from the client. Based on a partial knowledge of the network, we developed a scheduling algorithm that exploits the scalable video coding (SVC) properties to estimate packets importance and that takes into account packet delay dependencies to better anticipate congestion situations. Compared to more conventional streaming systems, experimental results show that our approach allows to better face network condition degradation like bandwidth reduction or packet error rate increase. Nicolas Tizon, Béatrice Pesquet-Popescu, Marco Cagnazzo |
ICIP | 3 |
| 2009 | Dense disparity estimation in multiview video codingabstractMultiview video coding is an emerging application where, in addition to classical temporal prediction, an efficient disparity prediction should be performed in order to achieve the best compression performance. A popular coder is the multiview video coding (MVC) extension of H.264/AVC, which uses a block-based disparity estimation (just like temporal prediction in H.264/AVC). In this paper, we propose to improve the MVC extension by using a dense estimation method that generates a smooth disparity map with ideally infinite precision. The obtained disparity is then segmented and efficiently encoded by using a rate-distortion optimization technique. Experimental results show that significant gains can be obtained compared to the block-based disparity estimation technique used in the MVC extension. Ismaël Daribo, Mounir Kaaniche, Wided Miled, Marco Cagnazzo, Béatrice Pesquet-Popescu |
MMSP | 4 |
| 2009 | Estimation of quantization noise for adaptive-prediction lifting schemesabstractThe lifting scheme represents an easy way of implementing the wavelet transform and of constructing new content-adapted transforms. However, the adaptive version of lifting schemes can result in strongly non-isometric transforms. This can be a major limitation, since all most successful coding techniques rely on the distortion estimation in the transform domain. In this paper we focus on the problem of evaluating the reconstruction distortion (due to quantization noise) in the wavelet domain when a non-isometric adaptive-prediction lifting scheme is used. The problem arises since these transforms are nonlinear, and so common techniques for distortion evaluation cannot be used in this case. We circumvent the difficulty by computing an equivalent time-varying linear filter, for which it is possible to generalize the distortion computation technique. In addition to the theoretical formulation of the distortion estimation, in this paper we provide experimental results proving the reliability of this estimation, and the consequent improvement of RD performance, thanks to a more effective resource allocation which can be performed in the transform domain. Sara Parrilli, Marco Cagnazzo, Béatrice Pesquet-Popescu |
MMSP | 2 |
| 2009 | Improving H.264 performances by quantization of motion vectorsabstractThe coding resources used for motion vectors (MVs) can attain quite high ratios even in the case of efficient video coders like H.264, and this can easily lead to suboptimal rate-distortion performance. In a previous paper, we proposed a new coding mode for H.264 based on the quantization of motion vectors (QMV). We only considered the case of 16 times 16 partitions for motion estimation and compensation. That method allowed us to obtain an improved trade-off in the resource allocation between vectors and coefficients, and to achieve better rate-distortion performances with respect to H.264. In this paper, we build on the proposed QMV coding mode, extending it to the case of macroblock partition into smaller blocks. This issue requires solving some problems mainly related to the motion vector coding. We show how this task can be performed efficiently in our framework, obtaining further improvements over the standard coding technique. Silvia Corrado, Marie Andrée Agostini, Marco Cagnazzo, Marc Antonini, Guillaume Laroche, Joël Jung |
PCS | 3 |
| 2008 | Distortion evaluation in transform domain for adaptive lifting schemesabstractIn this paper we study the problem of evaluating the reconstruction distortion in the wavelet domain when adaptive lifting schemes (ALS) are used for the direct and inverse transform. The distortion evaluation is necessary in order to perform efficient resource allocation over the transform coefficients. ALS is a non-linear transformation, which prevents using common techniques for distortion evaluation. However we show the equivalence of this non-linear scheme with a time-varying linear filter, and we generalize the distortion computation technique to it. Experiments show that the proposed method allows a reliable estimation of the distortion in the transform domain. This results in improved coding performance. Sara Parrilli, Marco Cagnazzo, Béatrice Pesquet-Popescu |
MMSP | 2 |
| 2007 | Improved Class-Based Coding of Multispectral Images With Shape-Adaptive Wavelet TransformabstractIn this letter, we improve the class-based transform-coding scheme proposed by Gelli and Poggi for the compression of multispectral images. The original spatial-coding tools, 1-D discrete cosine transform and scalar quantization, are replaced by shape-adaptive wavelet transform and set partitioning in hierarchical trees. Numerical experiments show that the improved technique outperforms the original one for medium- to high-quality compression and is consistently superior to all reference techniques. Marco Cagnazzo, Sara Parrilli, Giovanni Poggi, Luisa Verdoliva |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2007 | Optimal Motion Estimation for Wavelet Motion Compensated Video CodingabstractWavelet-based coding is emerging as a promising framework for efficient and scalable compression of video. Nevertheless, a number of basic tools currently employed in this field have been conceived for hybrid block-based transform coding. This is the case of motion estimation, which generally aims to minimize the energy or the absolute sum of prediction error. However, as wavelet video coders do not employ predictive coding, this is no longer an optimal approach. In this paper we study the problem of the theoretical optimal criterion for wavelet-based video coders, using coding gain as merit figure. A simple solution has been found for a peculiar but useful class of temporal filters. Experiments confirm that the optimally estimated vectors increase the coding gain as well as the performance of a complete video coder, but at the cost of an augmented complexity. Marco Cagnazzo, F. Castaldo, Thomas André, Marc Antonini, Michel Barlaud |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Region-Based Transform Coding of Multispectral ImagesabstractWe propose a new efficient region-based scheme for the compression of multispectral remote-sensing images. The region-based description of an image comprises a segmentation map, which singles out the relevant regions and provides their main features, followed by the detailed (possibly lossless) description of each region. The map conveys information on the image structure and could even be the only item of interest for the user; moreover, it enables the user to perform a selective download of the regions of interest, or can be used for high-level data mining and retrieval applications. This approach, with the multiple pieces of information required, may seem inherently inefficient. The goal of this research is to show that, by carefully selecting the appropriate segmentation and coding tools, region-based compression of multispectral images can be also effective in a rate-distortion sense, thus providing an image description that is both insightful and efficient. To this end, we define a generic coding scheme, based on Bayesian image segmentation and on transform coding, where several key design choices, however, are left open for optimization, from the type of transform, to the rate allocation procedure, and so on. Then, through an extensive experimental phase on real-world multispectral images, we gain insight on such key choices, and finally single out an efficient and robust coding scheme, with Bayesian segmentation, class-adaptive Karhunen-Loève spectral transform, and shape-adaptive wavelet spatial transform, which outperforms state-of-the-art and carefully tuned conventional techniques, such as JPEG-2000 multicomponent or SPIHT-based coders. Marco Cagnazzo, Giovanni Poggi, Luisa Verdoliva |
IEEE Trans. Image Process. | 1 |
| 2006 | Adaptive Region-Based Compression of Multispectral ImagesabstractThe region-based description of multispectral images enables important high-level tasks such as data mining and retrieval, and region-of-interest selection. In order to obtain an efficient representation of such images we resort to adaptive transform coding techniques. Such techniques, however, require a considerable information overhead, which must be carefully managed to obtain a satisfactory rate-distortion performance. In this work we develop several region-based coding schemes and compare them with conventional (non-adaptive) and class-based schemes, so as to single out the rate-distortion gains/losses of this approach. Marco Cagnazzo, Raffaele Gaetano, Sara Parrilli, Luisa Verdoliva |
ICIP | 1 |
| 2006 | Trading off quality and complexity for a HVQ-based video codec on portable devices
Marco Cagnazzo, Francesco Delfino, Luca Vollero, Andrea Zinicola |
J. Vis. Commun. Image Represent. | 1 |
| 2006 | Low-complexity compression of multispectral images based on classified transform coding
Marco Cagnazzo, Luca Cicala, Giovanni Poggi, Luisa Verdoliva |
Signal Process. Image Commun. | 1 |
| 2005 | A comparison of flat and object-based transform coding techniques for the compression of multispectral imagesabstractIn this work we implement and compare several state-of-the-art transform coding schemes for the compression of multispectral images, in order to better understand which elements have a deeper impact on the overall performance, and which tools guarantee the best results. All schemes are based on Karhunen-Loeve transform and/or wavelet transform, in various combinations, and use SPIHT as the coding engine. Moreover, besides the ordinary techniques, their object-based counterparts are also examined, so as to study the viability of such approach [M. Cagnazzo et al., Oct 2004] for these images. Whenever possible, an optimal rate allocation strategy is applied. The experiments, performed on images acquired by two different sensors, highlight the superiority of KLT as spectral transform; the rough equivalence between object-based and ordinary techniques in terms of rate-distortion performance; and the importance of the optimal allocation. Marco Cagnazzo, Giovanni Poggi, Luisa Verdoliva |
ICIP (1) | 1 |
| 2005 | Costs and advantages of shape-adaptive wavelet transform for region-based image codingabstractRegion-based encoding techniques have been long investigated for the compression of still images and video sequences and have recently gained much popularity, as testified by the object-based nature of the MPEG-4 video coding standard. This work aims at analyzing costs and advantages of implementing such an approach by shape-adaptive wavelet transform and shape-adaptive SPIHT. The analysis of several performance measures in a number of experiments confirm the potential of wavelet-based region-based approach, and provide insight about what performance gains and losses can be expected in various operative conditions. Marco Cagnazzo, Giovanni Poggi, Luisa Verdoliva |
ICIP (3) | 1 |
| 2004 | (N, 0) motion-compensated lifting-based wavelet transformabstractMotion compensation has been widely used in both DCT- and wavelet-based video coders for years. The recent success of the temporal wavelet transform based on motion-compensated lifting suggests that a high-performance, scalable wavelet video coder may soon outperform the best DCT-based coders. However, motion-compensated lifting does not implement exactly its transversal equivalent unless certain conditions on motion are satisfied. We review those conditions, and we discuss their importance. We derive a new class of temporal transforms, the so-called 1-N transversal or (N,0) lifting transforms, that are particularly interesting if those conditions on motion are not satisfied. We compare experimentally the 1-3 and 5-3 motion compensated wavelet transforms for the ubiquitous block-motion model used in all video compression standards. For this model, the 1-3 transform outperforms the 5-3 transform due to the need to transmit additional motion information in the latter case. This interesting result, however, does not extend to motion models satisfying the transversal/lifting equivalence conditions. Thomas André, Marco Cagnazzo, Marc Antonini, Michel Barlaud, Nikola Bozinovic, Janusz Konrad |
ICASSP (3) | 2 |
| 2004 | A model-based motion compensated video coder with JPEG2000 compatibilityabstractWe present a highly scalable wavelet-based video coder, featuring a scan-based motion-compensated temporal wavelet transform (WT) with lifting schemes which have been specially designed for video. Output bitstream is compatible with JPEG2000, as it is used to compress temporal subbands (SBs). Rate allocation among SBs is done by means of an optimal algorithm, which requires SBs rate-distortion (RD) curves. We propose a model-based approach allowing us to compute these curves with a considerable reduction in complexity. The use of temporal WT and JPEG2000 guarantees high scalability. Marco Cagnazzo, Thomas André, Marc Antonini, Michel Barlaud |
ICIP | 1 |
| 2004 | Region-oriented compression of multispfctral images by shape-adaptive wavelet transform and sphitabstractWe present a new technique for the compression of remote-sensing hyperspectral images based on wavelet transform and zerotree coding of coefficients. In order to improve encoding efficiency, the image is first segmented in a small number of regions with homogeneous texture. Then, a shape-adaptive wavelet transform is carried out on each region and the resulting coefficients are finally encoded by a shape-adaptive version of SPIHT. Thanks to the segmentation map (sent as a side information) region boundaries are faithfully preserved and selective encoding strategies can be easily implemented. In addition, by-now homogeneous region textures can be more efficiently encoded. Marco Cagnazzo, Giovanni Poggi, Luisa Verdoliva, Andrea Zinicola |
ICIP | 1 |
| 2004 | Compression of multitemporal remote sensing images through Bayesian segmentationabstractMultitemporal remote sensing images are useful tools for many applications in natural resource management. Compression of this kind of data is an issue of interest, yet, only a few paper address it specifically, while general-purpose compression algorithms are not well suited to the problem, as they do not exploit the strong correlation among images of a multitemporal set of data. Here we propose a coding architecture for multitemporal images, which takes advantage of segmentation in order to compress data. Segmentation subdivides images into homogeneous regions, which can be efficiently and independently encoded. Moreover this architecture provides the user with a great flexibility in transmitting and retrieving only data of interest Marco Cagnazzo, Giovanni Poggi, Giuseppe Scarpa, Luisa Verdoliva |
IGARSS | 1 |
| 2004 | A smoothly scalable and fully JPEG2000-compatible video coderabstractIn this paper, we analyze the scalability properties of the JPEG2000-compatible video encoder presented in M. Cagnazzo et al., (2004), and we improve its performances by presenting a new technique for an efficient motion vectors (MVs) encoding, producing a motion bitstream also compatible with JPEG2000. Our study shows that, thanks to our encoding strategy and to our peculiar temporal filters, scalably encoded sequences have the same or almost the same quality than non-scalably encoded ones: this is what we call smooth scalability. We also compared our encoder performances with the recent H.264 standard, showing comparable or sometimes better performances. Marco Cagnazzo, Thomas André, Marc Antonini, Michel Barlaud |
MMSP | 1 |
| 2002 | The advantage of segmentation in SAR image compressionabstractSAR images are severely degraded by speckle, and filtering is therefore a common practice. Filtering is especially useful before compression, to avoid spending valuable resources to represent noise; unfortunately, it also degrades important image features, like region boundaries. To overcome this problem, one can resort to a segmentation-based compression scheme, which allows one to preserve region boundaries, carry out intense denoising, and improve overall performance. In this work we assess the potential of segmentation-based compression through controlled experiments on synthetic SAR images. Numerical results seem to confirm the validity of this approach. Marco Cagnazzo, Giovanni Poggi, Luisa Verdoliva |
IGARSS | 1 |