VLDB 2026 Research / reviewers in the wild / expert
Thomas Maugey
dblp:53/7528
· DBLP profile ↗
75ranked-venue papers
19as first author
20since 2021 · last 2025
0000-0002-7149-0823ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 74 · 19 first-author · 20 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OSLO-IC: On-the-Sphere Learned Omnidirectional Image Compression with Attention Modules and Spatial ContextabstractDeveloping effective 360-degree (spherical) image compression techniques is crucial for technologies like virtual reality and automated driving. This paper advances the state-of-the-art in on-the-sphere learning (OSLO) for omnidirectional image compression framework by proposing spherical attention modules, residual blocks, and a spatial autoregressive context model. These improvements achieve a 23.1% bit rate reduction in terms of WS-PSNR BD rate. Additionally, we introduce a spherical transposed convolution operator for upsampling, which reduces trainable parameters by a factor of four compared to the pixel shuffling used in the OSLO framework, while maintaining similar compression performance. Therefore, in total, our proposed method offers significant rate savings with a smaller architecture and can be applied to any spherical convolutional application. Paul Wawerek-López, Navid Mahmoudian Bidgoli, Pascal Frossard, André Kaup, Thomas Maugey |
ICASSP | 5 |
| 2025 | Efficient Constraining of Transcoding in DNA-Based Image StorageabstractDNA has emerged as a promising alternative for long-term data storage due to its high capacity, durability, and low-energy potential. However, storing data in DNA presents several challenges. First, it requires complex and costly biochemical processes, making efficient compression crucial to reducing DNA synthesis time and cost. Second, these processes are prone to errors that must be avoided and/or corrected. In particular, homopolymers (repetitions of the same nucleotide) are a well-known source of errors during the sequencing step. Avoiding such repetitions helps mitigate errors but introduces a constraint that may increase the data compression rate. In this paper, we propose two transcoding methods that address these two key challenges: reducing data rate and minimizing errors. The first method strictly enforces the error-minimization constraint by eliminating homopolymers of a certain length, at the cost of an increased data rate. In contrast, the second method accepts a slight increase in homopolymers. However, we show that these increases remain limited (2.14% increase in compression rate for the first method and 0.39% homopolymer rate for the second). These two approaches demonstrate that it is possible to efficiently constrain transcoding while balancing error minimization and compression performance. Sara Al Sayyed, Aline Roumy, Thomas Maugey |
ICIP | 3 |
| 2025 | SCALED: Surrogate-gradient for Codec-Aware Learning of Downsampling in ABR StreamingabstractThe rapid growth in video consumption has introduced significant challenges to modern streaming architectures. Over-the-Top (OTT) video delivery now predominantly relies on Adaptive Bitrate (ABR) streaming, which dynamically adjusts bitrate and resolution based on client-side constraints such as display capabilities and network bandwidth. This pipeline typically involves downsampling the original high-resolution content, encoding and transmitting it, followed by decoding and upsampling on the client side. Traditionally, these processing stages have been optimized in isolation, leading to suboptimal end-to-end rate-distortion (R-D) performance. The advent of deep learning has spurred interest in jointly optimizing the ABR pipeline using learned resampling methods. However, training such systems end-to-end remains challenging due to the non-differentiable nature of standard video codecs, which obstructs gradient-based optimization. Recent works have addressed this issue using differentiable proxy models, based either on deep neural networks or hybrid coding schemes with differentiable components such as soft quantization, to approximate the codec behavior. While differentiable proxy codecs have enabled progress in compression-aware learning, they remain approximations that may not fully capture the behavior of standard, non-differentiable codecs. To our knowledge, there is no prior evidence demonstrating the inefficiencies of using standard codecs during training. In this work, we introduce a novel framework that enables end-to-end training with real, non-differentiable codecs by leveraging data-driven surrogate gradients derived from actual compression errors. It facilitates the alignment between training objectives and deployment performance. Experimental results show a 5.19\% improvement in BD-BR (PSNR) compared to codec-agnostic training approaches, consistently across the entire rate-distortion convex hull spanning multiple downsampling ratios. Esteban Pesnel, Julien Le Tanou, Michaël Ropert, Thomas Maugey, Aline Roumy |
PCS | 4 |
| 2025 | Compact image representation for content-based image retrieval in DNA data storage
Sara Al Sayyed, Aline Roumy, Thomas Maugey, Nicolas Lobato-Dauzier, Anthony J. Genot |
PCS | 3 |
| 2025 | Linearly Transformed Color Guide for Low-Bitrate Diffusion-Based Image CompressionabstractThis study addresses the challenge of controlling the global color aspect of images generated by a diffusion model without training or fine-tuning. We rewrite the guidance equations to ensure that the outputs are closer to a known color map, without compromising the quality of the generation. Our method results in new guidance equations. In the context of color guidance, we show that the scaling of the guidance should not decrease but rather increase throughout the diffusion process. In a second contribution, our guidance is applied in a compression framework, where we combine both semantic and general color information of the image to decode at very low cost. We show that our method is effective in improving the fidelity and realism of compressed images at extremely low bit rates ($10^{-2}$bpp), performing better on these criteria when compared to other classical or more semantically oriented approaches. The implementation of our method is available on gitlab athttps://gitlab.inria.fr/tbordin/color-guidance. Tom Bordin, Thomas Maugey |
IEEE Trans. Image Process. | 2 |
| 2024 | CoCliCo: Extremely Low Bitrate Image Compression Based on CLIP Semantic and Tiny Color MapabstractCoding algorithms are usually designed to pixel-wisely reconstruct images, which limits the expected gains in terms of compression. In this work, we introduce a semantic compressed representation for images: CoCliCo. We encode the inputs into a CLIP latent vector and a tiny color map, and we use a conditional diffusion model for reconstruction. When compared to the most recent traditional and generative coders, our approach reaches drastic compression gains while keeping most of the high-level information and a good level of realism. Tom Bachard, Tom Bordin, Thomas Maugey |
PCS | 3 |
| 2023 | Learning on Entropy Coded Images with CNNabstractWe propose an empirical study to see whether learning with convolutional neural networks (CNNs) on entropy coded data is possible. First, we define spatial and semantic closeness, two key properties that we experimentally show to be necessary to guarantee the efficiency of the convolution. Then, we show that these properties are not satisfied by the data processed by an entropy coder. Despite this, our experimental results show that learning in such difficult conditions is still possible, and that the performance are far from a random guess. These results have been obtained thanks to the construction of CNN architectures designed for 1D data (one based on VGG, the other on ResNet). Finally, we propose some experiments that explain why CNN are still performing reasonably well on entropy coded data. Rémi Piau, Thomas Maugey, Aline Roumy |
ICASSP | 2 |
| 2023 | Vanishing Point Aided Hash-Frequency Encoding for Neural Radiance Fields (NeRF) from Sparse 360°InputabstractNeural Radiance Fields (NeRF) enable novel view synthesis of 3D scenes when trained with a set of 2D images. One of the key components of NeRF is the input encoding, i.e. mapping the coordinates to higher dimensions to learn high-frequency details, which has been proven to increase the quality. Among various input mappings, hash encoding is gaining increasing attention for its efficiency. However, its performance on sparse inputs is limited. To address this limitation, we propose a new input encoding scheme that improves hash-based NeRF for sparse inputs, i.e. few and distant cameras, specifically for 360° view synthesis. In this paper, we combine frequency encoding and hash encoding and show that this combination can increase dramatically the quality of hash-based NeRF for sparse inputs. Additionally, we explore scene geometry by estimating vanishing points in omnidirectional images (ODI) of indoor and city scenes in order to align frequency encoding with scene structures. We demonstrate that our vanishing point-aided scene alignment further improves deterministic and non-deterministic encodings on image regression and NeRF tasks where sharper textures and more accurate geometry of scene structures can be reconstructed. Thomas Maugey, Sebastian Knorr, Christine Guillemot |
ISMAR | 2 |
| 2023 | Semantic Based Generative Compression of Images for Extremely Low BitratesabstractWe propose a framework for image compression in which the fidelity criterion is replaced by a semantic and quality preservation objective. Encoding the image thus becomes a simple extraction of semantic, enabling to reach drastic compression ratio. The decoding side is handled by a generative model relying on the diffusion process for the reconstruction of images. We first propose to describe the semantic using low resolution segmentation maps as guide. We further improve the generation, introducing colors map guidance without retraining the generative decoder. We show that it is possible to produce images of high visual quality with preserved semantic at extremely low bitrates when compared with classical codecs. Tom Bordin, Thomas Maugey |
MMSP | 2 |
| 2023 | Towards Digital Sobriety: Why Improving the Energy Efficiency of Video Streaming is Not EnoughabstractIPCC conclusions are unequivocal: we must divide our greenhouse gas emissions by two before 2030 if we want to maintain the global warming below 1.5°C in 2100. Hence, it becomes urgent to aim sobriety. Contrary to what is often claimed, digital technologies must also target global emission reduction, as their impact on the climate is huge and exploding every year. Among the digital world's emissions, those related to video processing and streaming are significant. At the same time, a lot of research efforts are currently done to reduce the energy consumed by video transmission algorithms or infrastructures. In this paper, we demonstrate that, even though such research works are crucial, they are not sufficient to enable global video streaming emissions reductions. The conclusion is that we must collectively think of other complementary solutions. Thomas Maugey |
MMSP | 1 |
| 2022 | Semantic Alignment for Multi-Item CompressionabstractCoding algorithms usually compress independently the images of a collection, in particular when the correlation between them only resides at the semantic level, i.e., information related to the high-level image content. In this work, we propose a coding solution able to exploit this semantic redundancy to decrease the storage cost of data collections. First we introduce the multi-item compression framework. Then we derive a loss term to shape the latent space of a variational auto-encoder so that the latent vectors of semantically identical images can be aligned. Finally, we experimentally demonstrate that this alignment leads to a more compact representation of the data collection. Tom Bachard, Anju Jose Tom, Thomas Maugey |
ICIP | 3 |
| 2022 | Omni-NeRF: Neural Radiance Field from 360° Image CapturesabstractThis paper tackles the problem of novel view synthesis (NVS) from 360° images with imperfect camera poses or intrinsic parameters. We propose a novel end-to-end framework for training Neural Radiance Field (NeRF) models given only 360° RGB images and their rough poses, which we refer to as Omni-NeRF. We extend the pinhole camera model of NeRF to a more general camera model that better fits omni-directional fish-eye lenses. The approach jointly learns the scene geometry and optimizes the camera parameters without knowing the fisheye projection. Thomas Maugey, Sebastian Knorr, Christine Guillemot |
ICME | 2 |
| 2022 | Motion Compensation-based Low-Complexity Decoder Side Depth Estimation for MPEG Immersive VideoabstractDecoder-Side Depth Estimation (DSDE) is a system firstly enabled in the novel MPEG Immersive Video (MIV) coding standard. In DSDE, only texture components are coded, while the depth is estimated at the decoder-side. This is motivated by previous work, which has shown high coding gain and pixel rate savings in DSDE. However, the computational complexity remains a concern, as high quality depth search has a high runtime and memory requirement. In this work we extend the concept of depth estimation to depth recovery. Using this mode, the decoder-side depth information is recovered through motion compensation utilizing the displacement vectors contained in the texture bitstream. This strategy enables us to replace most of the complex depth estimation processes with a simple motion compensation step, a decision that is drawn on the encoder-side and signaled per coding unit. With only minor losses in terms of synthesis PSNR and similar perceptual quality in terms of MS-SSIM, the complexity is significantly reduced. Depending on the acceptable loss, up to 80 % of the moving objects depth may be motion compensated instead of estimated by a depth estimator translating into a speed-up of a factor of 104 for inter-frames compared to the reference depth estimator. Patrick Garus, Félix Henry, Thomas Maugey, Christine Guillemot |
MMSP | 3 |
| 2022 | Decoder Side Multiplane Images using Geometry Assistance SEI for MPEG Immersive VideoabstractThe MPEG Immersive Video (MIV) standard enables a novel technology denoted as decoder side depth estimation (DSDE) by introducing a dedicated Geometry Absent profile. In DSDE only texture information is coded and the corresponding geometry is reconstructed on the decoder side. MIV further enables the coding of side-information useful to the geometry reconstruction, denoted as Geometry Assistance SEI message. An emerging format for immersive video are Multiplane Images, which is investigated for feasibility in coding systems due to their promising rendering quality with complex sequences. In this work, we show that MIV can be used to construct block-based Multiplane Images on the decoder-side and to enhance the view synthesis performance utilizing the Geometry Assistance SEI. In a complexity-aware setting using only 32 planes, up to 6 dB of quality improvement is achieved compared to the reference. Patrick Garus, Félix Henry, Thomas Maugey, Christine Guillemot |
MMSP | 3 |
| 2022 | Immersive Video Coding: Should Geometry Information Be Transmitted as Depth Maps?abstractImmersive video often refers to multiple views with texture and scene geometry information, from which different viewports can be synthesized on the client side. To design efficient immersive video coding solutions, it is desirable to minimize bitrate, pixel rate and complexity. We investigate whether the classical approach of sending the geometry of a scene as depth maps is appropriate to serve this purpose. Previous work shows that bypassing depth transmission entirely and estimating depth at the client side improves the synthesis performance while saving bitrate and pixel rate. In order to understand if the encoder side depth maps contain information that is beneficial to be transmitted, we first explore a hybrid approach which enables partial depth map transmission using a block-based RD-based decision in the depth coding process. This approach reveals that partial depth map transmission may improve the rendering performance but does not present a good compromise in terms of compression efficiency. This led us to address the remaining drawbacks of decoder side depth estimation: complexity and depth map inaccuracy. We propose a novel system that takes advantage of high quality depth maps at the server side by encoding them into lightweight features that support the depth estimator at the client side. These features allow reducing the amount of data that has to be handled during decoder side depth estimation by 88%, which significantly speeds up the cost computation and the energy minimization of the depth estimator. Furthermore, −46.0% and −37.9% average synthesis BD-Rate gains are achieved compared to the classical approach with depth maps estimated at the encoder. Patrick Garus, Félix Henry, Joël Jung, Thomas Maugey, Christine Guillemot |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | OSLO: On-the-Sphere Learning for Omnidirectional Images and Its Application to 360-Degree Image CompressionabstractState-of-the-art 2D image compression schemes rely on the power of convolutional neural networks (CNNs). Although CNNs offer promising perspectives for 2D image compression, extending such models to omnidirectional images is not straightforward. First, omnidirectional images have specific spatial and statistical properties that can not be fully captured by current CNN models. Second, basic mathematical operations composing a CNN architecture, e.g., translation and sampling, are not well-defined on the sphere. In this paper, we study the learning of representation models for omnidirectional images and propose to use the properties of HEALPix uniform sampling of the sphere to redefine the mathematical tools used in deep learning models for omnidirectional images. In particular, we: i) propose the definition of a new convolution operation on the sphere that keeps the high expressiveness and the low complexity of a classical 2D convolution; ii) adapt standard CNN techniques such as stride, iterative aggregation, and pixel shuffling to the spherical domain; and then iii) apply our new framework to the task of omnidirectional image compression. Our experiments show that our proposed on-the-sphere solution leads to a better compression gain that can save 13.7% of the bit rate compared to similar learned models applied to equirectangular images. Also, compared to learning models based on graph convolutional networks, our solution supports more expressive filters that can preserve high frequencies and provide a better perceptual quality of the compressed images. Such results demonstrate the efficiency of the proposed framework, which opens new research venues for other omnidirectional vision tasks to be effectively implemented on the sphere manifold. Navid Mahmoudian Bidgoli, Roberto Gerson De Albuquerque Azevedo, Thomas Maugey, Aline Roumy, Pascal Frossard |
IEEE Trans. Image Process. | 3 |
| 2021 | Rate-Distortion Optimized Motion Estimation for on-the-Sphere Compression of 360 VideosabstractOn-the-sphere compression of omnidirectional videos is a very promising approach. First, it saves computational complexity as it avoids to project the sphere onto a 2D map, as classically done. Second, and more importantly, it allows to achieve a better rate-distortion tradeoff, since neither the visual data nor its domain of definition are distorted. In this paper, the on-the-sphere compression [1] for omnidirectional still images is extended to videos. We first propose a complete review of existing spherical motion models. Then we pro-pose a new one called tangent-linear+t. We finally propose a rate-distortion optimized algorithm to locally choose the best motion model for efficient motion estimation/compensation. For that purpose, we additionally propose a finer search pattern, called spherical-uniform, for the motion parameters, which leads to a more accurate block prediction. The novel algorithm leads to rate-distortion gains compared to methods based on a unique motion model. Alban Marie, Navid Mahmoudian Bidgoli, Thomas Maugey, Aline Roumy |
ICASSP | 3 |
| 2021 | Evaluation Of Bitrate Ladders For Versatile Video CoderabstractMany video service providers take advantage of bitrate ladders in adaptive HTTP video streaming to account for different network states and user display specifications by providing bitrate/resolution pairs that best fit client's network conditions and display capabilities. These bitrate ladders, however, differ when using different codecs and thus the couples bitrate/resolution differ as well. In addition, bitrate ladders are based on previously available codecs (H.264/MPEG4-AVC, HEVC, etc.), i.e. codecs that are already in service, hence the introduction of new codecs e.g. Versatile Video Coding (VVC) requires re-analyzing these ladders. For that matter, we will analyze the evolution of the bitrate ladder when using VVC. We show how VVC impacts this ladder when compared to HEVC and H.264/AVC and in particular, that there is no need to switch to lower resolutions at the lower bitrates defined in the Call for Evidence on Transcoding for Network Distributed Video Coding (CfE). Reda Kaafarani, Médéric Blestel, Thomas Maugey, Michaël Ropert, Aline Roumy |
VCIP | 3 |
| 2021 | Rate-Distortion Optimized Graph Coarsening and Partitioning for Light Field CodingabstractGraph-based transforms are powerful tools for signal representation and energy compaction. However, their use for high dimensional signals such as light fields poses obvious problems of complexity. To overcome this difficulty, one can consider local graph transforms defined on supports of limited dimension, which may however not allow us to fully exploit long-term signal correlation. In this paper, we present methods to optimize local graph supports in a rate distortion sense for efficient light field compression. A large graph support can be well adapted for compression efficiency, however at the expense of high complexity. In this case, we use graph reduction techniques to make the graph transform feasible. We also consider spectral clustering to reduce the dimension of the graph supports while controlling both rate and complexity. We derive the distortion and rate models which are then used to guide the graph optimization. We describe a complete light field coding scheme based on the proposed graph optimization tools. Experimental results show rate-distortion performance gains compared to the use of fixed graph support. The method also provides competitive results when compared against HEVC-based and the JPEG Pleno light field coding schemes. We also assess the method against a homography-based low rank approximation and a Fourier disparity layer based coding method. Mira Rizkallah, Thomas Maugey, Christine Guillemot |
IEEE Trans. Image Process. | 2 |
| 2021 | Fine Granularity Access in Interactive Compression of 360-Degree Images Based on Rate-adaptive Channel CodesabstractIn this paper, we propose a new interactive compression scheme for omnidirectional images. This requires two characteristics: efficient compression of data, to lower the storage cost, and random access ability to extract part of the compressed stream requested by the user (for reducing the transmission rate). For efficient compression, data needs to be predicted by a series of references that have been pre-defined and compressed. This contrasts with the spirit of random accessibility. We propose a solution for this problem based on incremental codes implemented by rate-adaptive channel codes. This scheme encodes the image while adapting to any user request and leads to an efficient coding that is flexible in extracting data depending on the available information at the decoder. Therefore, only the information that is needed to be displayed at the user's side is transmitted during the user's request, as if the request was already known at the encoder. The experimental results demonstrate that our coder obtains a better transmission rate than the state-of-the-art tile-based methods at a small cost in storage. Moreover, the transmission rate grows gradually with the size of the request and avoids a staircase effect, which shows the perfect suitability of our coder for interactive transmission. Navid Mahmoudian Bidgoli, Thomas Maugey, Aline Roumy |
IEEE Trans. Multim. | 2 |
| 2020 | Sphere Mapping for Feature Extraction From 360° Fish-Eye CapturesabstractEquirectangular projection is commonly used to map 360° captures into planar representation, so that existent processing methods can be directly applied to such content. Such format introduces stitching distortions that could impact the efficiency of further processing such as camera pose estimation, 3D point localization and depth estimation. Indeed, even if some algorithms, mainly feature descriptors, tend to remap the projected images into a sphere, important radial distortions remain existent in the processed data. In this paper, we propose to adapt the spherical model to the geometry of the 360° fish-eye camera, and avoid the stitching process. We consider the angular coordinates of feature points on the sphere for evaluation. We assess the precision of different operations such as camera rotation angle estimation and 3D point depth calculation on spherical camera images. Experimental results show that the proposed fish-eye adapted sphere mapping allows more stability in angle estimation, as well as in 3D point localization, compared to the one on projected and stitched contents. Fatma Hawary, Thomas Maugey, Christine Guillemot |
MMSP | 2 |
| 2020 | Large Database Compression Based on Perceived InformationabstractLossy compression algorithms trade bits for quality, aiming at reducing as much as possible the bitrate needed to represent the original source (or set of sources), while preserving the source quality. In this letter, we propose a novel paradigm of compression algorithms, aimed at minimizing the information loss perceived by the final user instead of the actual source quality loss, under compression rate constraints. As main contributions, we first introduce the concept of perceived information (PI), which reflects the information perceived by a given user experiencing a data collection, and which is evaluated as the volume spanned by the sources features in a personalized latent space. We then formalize the rate-PI optimization problem and propose an algorithm to solve this compression problem. Finally, we validate our algorithm against benchmark solutions with simulation results, showing the gain in taking into account users' preferences while also maximizing the perceived information in the feature domain. Thomas Maugey, Laura Toni |
IEEE Signal Process. Lett. | 1 |
| 2020 | Optimal Reference Selection for Random Access in Predictive Coding SchemesabstractData acquired over long periods of time like High Definition (HD) videos or records from a sensor over long time intervals, have to be efficiently compressed, to reduce their size. The compression has also to allow efficient access to random parts of the data upon request from the users. Efficient compression is usually achieved with prediction between data points at successive time instants. However, this creates dependencies between the compressed representations, which is contrary to the idea of random access. Prediction methods rely in particular on reference data points, used to predict other data points. The placement of these references balances compression efficiency and random access. Existing solutions to position the references use ad hoc methods. In this paper, we study this joint problem of compression efficiency and random access. We introduce the storage cost as a measure of the compression efficiency and the transmission cost for the random access ability. We express the reference placement problem that trades storage with transmission cost as an integer linear programming problem. Considering additional assumptions on the sources and coding methods reduces the complexity of the search space of the optimization problem. Moreover, we show that the classical periodic placement of the references is optimal, when the encoding costs of each data point are equal and when requests of successive data points are made. In this particular case, a closed-form expression of the optimal period is derived. Finally, the proposed optimal placement strategy is compared with an ad hoc method, where the references correspond to sources where the prediction does not help reducing significantly the encoding cost. The proposed optimal algorithm shows a bit saving of -20% with respect to the ad hoc method. Mai Quyen Pham, Aline Roumy, Thomas Maugey, Elsa Dupraz, Michel Kieffer |
IEEE Trans. Commun. | 3 |
| 2020 | Geometry-Aware Graph Transforms for Light Field Compact RepresentationabstractThe paper addresses the problem of energy compaction of dense 4D light fields by designing geometry-aware local graph-based transforms. Local graphs are constructed on super-rays that can be seen as a grouping of spatially and geometry-dependent angularly correlated pixels. Both non separable and separable transforms are considered. Despite the local support of limited size defined by the super-rays, the Laplacian matrix of the non separable graph remains of high dimension and its diagonalization to compute the transform eigen vectors remains computationally expensive. To solve this problem, we then perform the local spatio-angular transform in a separable manner. We show that when the shape of corresponding super-pixels in the different views is not isometric, the basis functions of the spatial transforms are not coherent, resulting in decreased correlation between spatial transform coefficients. We hence propose a novel transform optimization method that aims at preserving angular correlation even when the shapes of the super-pixels are not isometric. Experimental results show the benefit of the approach in terms of energy compaction. A coding scheme is also described to assess the rate-distortion perfomances of the proposed transforms and is compared to state of the art encoders namely HEVC-lozenge [1], JPEG pleno 1.1 [2], HEVC-pseudo [3] and HLRA [4]. Mira Rizkallah, Xin Su 0003, Thomas Maugey, Christine Guillemot |
IEEE Trans. Image Process. | 3 |
| 2020 | Prediction and Sampling With Local Graph Transforms for Quasi-Lossless Light Field CompressionabstractGraph-based transforms have been shown to be powerful tools in terms of image energy compaction. However, when the size of the support increases to best capture signal dependencies, the computation of the basis functions becomes rapidly untractable. This problem is in particular compelling for high dimensional imaging data such as light fields. The use of local transforms with limited supports is a way to cope with this computational difficulty. Unfortunately, the locality of the support may not allow us to fully exploit long term signal dependencies present in both the spatial and angular dimensions of light fields. This paper describes sampling and prediction schemes with local graph-based transforms enabling to efficiently compact the signal energy and exploit dependencies beyond the local graph support. The proposed approach is investigated and is shown to be very efficient in the context of spatio-angular transforms for quasi-lossless compression of light fields. Mira Rizkallah, Thomas Maugey, Christine Guillemot |
IEEE Trans. Image Process. | 2 |
| 2019 | Graph-Based Spatio-Angular Prediction for Quasi-Lossless Compression of Light FieldsabstractGraph-based transforms have been shown to be powerful tools for image compression. However, the computation of the basis functions becomes rapidly untractable when the support increases, i.e. when the dimension of the data is high as in the case of light fields. Local transforms with limited supports have been investigated to cope with this difficulty. Nevertheless, the locality of the support may not allow us to fully exploit long term dependencies in the signal. In this paper, we describe a graph based prediction solution that allows taking advantage of intra prediction mechanisms as well as of the good energy compaction properties of the graph transform. The approach relies on a separable spatio-angular transform and derives low frequency spatio-angular coefficients from one single compressed reference view and from the high angular frequency coefficients. In the tests, we used HEVC-Intra, with QP=0, to encode the reference frame with high quality. The high angular frequency coefficients containing very little energy are coded using a simple entropy coder. The approach is shown to be very efficient in a context of high quality quasi-lossless compression of light fields. Mira Rizkallah, Thomas Maugey, Christine Guillemot |
DCC | 2 |
| 2019 | A Geometry-aware Framework for Compressing 3D Mesh TexturesabstractThis paper proposes a novel prediction tool for improving the compression performance of texture atlases. This algorithm, called Geometry-Aware (GA) intra coding, takes advantage of the topology of the associated 3D meshes, in order to reduce the redundancies in the texture map. For texture processing, the concept of the conventional intra prediction, used in video compression, has been adapted to consider neighboring information on the 3D surface. We have also studied how this prediction tool can be integrated into a complete coding solution. In particular, a block scanning strategy and a graph-based transform for residual coding have been proposed. Results show that the knowledge of the mesh topology significantly improves the compression efficiency of texture atlases. Fatemeh Nasiri, Navid Mahmoudian Bidgoli, Frédéric Payan, Thomas Maugey |
ICASSP | 4 |
| 2019 | Evaluation Framework for 360-Degree Visual Content Compression with User View-Dependent TransmissionabstractImmersive visual experience can be obtained by allowing the user to navigate in a 360-degree visual content. These contents are stored in high resolution and need a lot of space on the server to store them. The transmission depends on the user's request and only the spatial region which is requested by the user is transmitted to avoid wasting network bandwidth. Therefore, storage and transmission rates are both critical. Splitting the rates into storage and transmission has not been formally considered in the literature for evaluating 360-degree content compression algorithms. In this paper, we propose a framework to evaluate the coding efficiency of 360-degree content while discriminating between storage and transmission rate and taking into account user dependency. This brings the flexibility to compare different coding methods based on the storage capacity on the server and network bandwidth of users. Navid Mahmoudian Bidgoli, Thomas Maugey, Aline Roumy |
ICIP | 2 |
| 2019 | FTV360: a multiview 360° video dataset with calibration parametersabstractIn this paper, we present a new dataset in order to serve as a support for researches in Free Viewpoint Television (FTV) and 6 degrees-of-freedom (6DoF) immersive communication. This dataset relies on a novel acquisition procedure consisting in a synchronized capture of a scene by 40 omnidirectional cameras. We have also developed a calibration solution that estimates the position and orientation of each camera with respect to a same reference. This solution relies on a regular calibration of each individual camera, and a graph-based synchronization of all these parameters. These videos and the calibration solution are made publicly available. Thomas Maugey, Laurent Guillo, Cedric Le Cam |
MMSys | 1 |
| 2019 | Intra-coding of 360-degree images on the sphereabstractOmni-directional images are characterized by their high resolution (usually 8K) and therefore require high compression efficiency. Existing methods project the spherical content onto one or multiple planes and process the mapped content with classical 2D video coding algorithms. However, this projection induces sub-optimality. Indeed, after projection, the statistical properties of the pixels are modified, the connectivity between neighboring pixels on the sphere might be lost, and finally, the sampling is not uniform. Therefore, we propose to process uniformly distributed pixels directly on the sphere to achieve high compression efficiency. In particular, a scanning order and a prediction scheme are proposed to exploit, directly on the sphere, the statistical dependencies between the pixels. A Graph Fourier Transform is also applied to exploit local dependencies while taking into account the 3D geometry. Experimental results demonstrate that the proposed method provides up to 5.6% bitrate reduction and on average around 2% bitrate reduction over state-of-the-art methods. Navid Mahmoudian Bidgoli, Thomas Maugey, Aline Roumy |
PCS | 2 |
| 2019 | A geometry-aware compression of 3D mesh texture with random accessabstractA 3D mesh object is usually represented as a combination of several entities including geometrical information (i.e., the triangles and their position in space) and a texture atlas/map (i.e. a giant 2D image containing all the texture information that is mapped to the 3D object at the rendering stage). This atlas is usually compressed using a conventional 2D image coder, thus without taking into account the geometrical information. Moreover, the whole image is usually decoded even though only a subpart of the mesh is observed by a user. In this paper, we propose a novel approach to compress a texture atlas of a 3D model that enables random access during decoding, and nevertheless takes into account the correlation driven by the geometrical information. The experimental results demonstrate the benefits of the proposed coder. Navid Mahmoudian Bidgoli, Thomas Maugey, Aline Roumy, Fatemeh Nasiri, Frédéric Payan |
PCS | 2 |
| 2019 | Bypassing Depth Maps Transmission For Immersive Video CodingabstractThis paper addresses several downsides of the system under development in MPEG-I for coding and transmission of immersive media. We present a solution, which enables Depth-Image-Based Rendering for immersive video applications, while lifting the requirement of transmitting depth information. Instead, we estimate the depth information on the client-side from the transmitted views. The approach leads to an impressive rate saving (37.3% in average). Preserving perceptual quality in terms of MS-SSIM of synthesized views, it yields to 24.6% rate reduction for the same quality of reconstructed views after residue transmission under the MPEG-I common test conditions. Simultaneously, the required pixel rate, i.e. the number of pixels processed per second by the decoder, is reduced by 50% for any test sequence. To the author's knowledge, this is the first time that such an approach is under consideration in the context of immersive video coding. Patrick Garus, Joël Jung, Thomas Maugey, Christine Guillemot |
PCS | 3 |
| 2018 | Rate-Distortion Performance of Sequential Massive Random Access to Gaussian Sources with MemoryabstractIn Sequential Massive Random Access (SMRA) [1, 2], a set of correlated sources is jointly encoded and stored on a server, and clients want to access to only a subset of the sources. Since the number of simultaneous clients can be huge, the server is only authorized to extract a bitstream from the stored data: no re-encoding can be performed before the transmission of a request. In this paper, we investigate the SMRA performance of lossy source coding of Gaussian sources with memory. In practical applications such as Free Viewpoint Television, this model permits to take into account not only inter but also intra correlation between sources. For this model, we provide the storage and transmission rates that are achievable for SMRA under some distortion constraint, and we consider two particular examples of Gaussian sources with memory. Elsa Dupraz, Thomas Maugey, Aline Roumy, Michel Kieffer |
DCC | 2 |
| 2018 | Graph-based Transforms for Predictive Light Field Compression based on Super-PixelsabstractIn this paper, we explore the use of graph-based transforms to capture correlation in light fields. We consider a scheme in which view synthesis is used as a first step to exploit inter-view correlation. Local graph-based transforms (GT) are then considered for energy compaction of the residue signals. The structure of the local graphs is derived from a coherent super-pixel over-segmentation of the different views. The GT is computed and applied in a separable manner with a first spatial unweighted transform followed by an inter-view GT. For the inter-view GT, both unweighted and weighted GT have been considered. The use of separable instead of non separable transforms allows us to limit the complexity inherent to the computation of the basis functions. A dedicated simple coding scheme is then described for the proposed GT based light field decomposition. Experimental results show a significant improvement with our method compared to the CNN view synthesis method and to the HEVC direct coding of the light field views. Mira Rizkallah, Xin Su 0003, Thomas Maugey, Christine Guillemot |
ICASSP | 3 |
| 2018 | Optimized Data Representation for Interactive Multiview NavigationabstractIn contrary to traditional media streaming services where a unique media content is delivered to different users, interactive multiview navigation applications enable users to choose their own viewpoints and freely navigate in a three-dimensional scene. The interactivity brings new challenges in addition to the classical rate-distortion tradeoff, which considers only the compression performance and viewing quality. On one hand, interactivity necessitates sufficient viewpoints for richer navigation; on the other hand, it requires to provide low bandwidth and delay costs for smooth navigation during view transitions. In this paper, we formally describe the novel tradeoffs posed by the navigation interactivity and classical rate-distortion criterion. Based on an original formulation, we look for the optimal design of the data representation by introducing novel rate and distortion models and practical solving algorithms. Experiments show that the proposed data representation method outperforms the baseline solution by providing lower resource consumptions and higher visual quality in all navigation configurations, which certainly confirms the potential of the proposed data representation in practical interactive navigation systems. Rui Ma 0006, Thomas Maugey, Pascal Frossard |
IEEE Trans. Multim. | 2 |
| 2017 | Correlation model selection for interactive video communicationabstractInteractive video communication has been recently proposed for multi-view videos. In this scheme, the server has to store the views as compact as possible, while being able to transmit them independently to the users, who are allowed to navigate interactively among the views, hence requesting a subset of them. To achieve this goal, the compression must be done using a model-based coding in which the correlation between the predicted view generated on the user side and the original view has to be modeled by a statistical distribution. In this paper we propose a framework for lossless fixed-length source coding to select a model among a candidate set of models that incurs the lowest extra rate cost to the system. Moreover, in cases where the depth image is available, we provide a method to estimate the correlation model. Navid Mahmoudian Bidgoli, Thomas Maugey, Aline Roumy |
ICIP | 2 |
| 2017 | Graph-based light fields representation and coding using geometry informationabstractThis paper describes a graph-based coding scheme for light fields (LF). It first adapts graph-based representations (GBR) to describe color and geometry information of LF. Graph connections describing scene geometry capture inter-view dependencies. They are used as the support of a weighted Graph Fourier Transform (wGFT) to encode disoccluded pixels. The quality of the LF reconstructed from the graph is enhanced by adding extra color information to the representation for a sub-set of sub-aperture images. Experiments show that the proposed scheme yields rate-distortion gains compared with HEVC based compression (directly compressing the LF as a video sequence by HEVC). Xin Su 0003, Mira Rizkallah, Thomas Maugey, Christine Guillemot |
ICIP | 3 |
| 2017 | Saliency-based navigation in omnidirectional imageabstractOmnidirectional images describe the color information at a given position from all directions. Affordable 360° cameras have recently been developed leading to an explosion of the 360° data shared on social networks. However, an omnidirectional image does not contain interesting content everywhere. Some part of the images are indeed more likely to be looked at by some users than others. Knowing these regions of interest might be useful for 360° image compression, streaming, retargeting or even editing. In this paper, we aim at modelling the user navigation within a 360° image, and detecting which parts of an omnidirectional content might draw users' attention. In particular, the paper proposes to aggregate and analyze 2D saliency detectors in different map projections, and also proposes a smooth navigation through the image to maximize saliency. Thomas Maugey, Olivier Le Meur, Zhi Liu 0003 |
MMSP | 1 |
| 2017 | Rate-Distortion Optimized Graph-Based Representation for Multiview Images With Complex Camera ConfigurationsabstractGraph-based representation (GBR) has recently been proposed for describing color and geometry of multiview video content. The graph vertices represent the color information, while the edges represent the geometry information, i.e., the disparity, by connecting corresponding pixels in two camera views. In this paper, we generalize the GBR to multiview images with complex camera configurations. Compared with the existing GBR, the proposed representation can handle not only horizontal displacements of the cameras but also forward/backward translations, rotations, etc. However, contrary to the usual disparity that is a 2-D vector (denoting horizontal and vertical displacements), each edge in GBR is represented by a 1-D disparity. This quantity can be seen as the disparity along an epipolar segment. In order to have a sparse (i.e., easy to code) graph structure, we propose a rate-distortion model to select the most meaningful edges. Hence the graph is constructed with "just enough" information for rendering the given predicted view. The experiments show that the proposed GBR allows high reconstruction quality with lower or equivalent coding rate than traditional depth-based representations. Xin Su 0003, Thomas Maugey, Christine Guillemot |
IEEE Trans. Image Process. | 2 |
| 2017 | Wide-Baseline Foreground Object Interpolation Using Silhouette Shape PriorabstractWe consider the synthesis of intermediate views of an object captured by two widely spaced and calibrated cameras. This problem is challenging because foreshortening effects and occlusions induce significant differences between the reference images when the cameras are far apart. That makes the association or disappearance/appearance of their pixels difficult to estimate. Our main contribution lies in disambiguating this ill-posed problem by making the interpolated views consistent with a plausible transformation of the object silhouette between the reference views. This plausible transformation is derived from an object-specific prior that consists of a nonlinear shape manifold learned from multiple previous observations of this object by the two reference cameras. The prior is used to estimate the evolution of the epipolar silhouette segments between the reference views. This information directly supports the definition of epipolar silhouette segments in the intermediate views, as well as the synthesis of textures in those segments. It permits to reconstruct the epipolar plane images (EPIs) and the continuum of views associated with the EPI volume, obtained by aggregating the EPIs. Experiments on synthetic and natural images show that our method preserves the object topology in intermediate views and deals effectively with the self-occluded regions and the severe foreshortening effect associated with wide-baseline camera configurations. Cédric Verleysen, Thomas Maugey, Pascal Frossard, Christophe De Vleeschouwer |
IEEE Trans. Image Process. | 2 |
| 2016 | Graph-based representation for multiview images with complex camera configurationsabstractInstead of lossily coding depth images resulting in undesirable geometric distortion, graph-based representation (GBR) describes disparity information as a graph with a controllable accuracy. In this paper, we propose a more compact graphical representation called GBR-plus to code both disparity and color information of a target view given a reference view. Specifically, first we differentiate between disocclusion holes (occluded spatial regions in the reference view) and rounding holes (insufficiently sampled regions in the reference view) in the synthesized target view, so that the decoder can optionally complete rounding holes via signal interpolation without coding overhead. Second, we use a compact graphical representation to delimit disparity-shifted boundaries of objects in the target view, which is coded losslessly. Finally, color pixels in disocclusion holes are predicted using adjacent background pixels as predictors, and prediction residuals in a local neighborhood are coded using Graph Fourier Transform (GFT). Experimental results show that GBR-plus outperforms previous GBR, and has comparable performance as HEVC at mid to high bitrates with lower encoder complexity. Xin Su 0003, Thomas Maugey, Christine Guillemot |
ICIP | 2 |
| 2016 | Temporal and Inter-View Consistent Error Concealment Technique for Multiview Plus Depth VideoabstractMultiview plus depth (MVD) is an emerging video format with many applications, including 3-D television and free viewpoint television. During the broadcast of a compressed MVD video, transmission errors may cause the loss of whole frames, resulting in significant degradation of video quality. Error concealment techniques have been widely used to deal with transmission errors in video communication. However, the existing solutions do not address the requirement that the reconstructed frames should be consistent with neighboring frames, i.e., corresponding pixels should have consistent color information. We propose a new consistency model for error concealment of MVD video that allows one to maintain a high level of consistency between frames of the same view (temporal consistency) and those of neighboring views (inter-view consistency). We then propose an algorithm that uses our model to implement concealment in a consistent way. Simulations with the reference software for the multiview video coding project of the joint video team of the ISO/IEC MPEG and ITU-T VCEG show that our method outperforms benchmark techniques, including a baseline approach based on the boundary matching algorithm, with respect to both reconstruction quality and view consistency. Shadan Khan Khattak, Thomas Maugey, Raouf Hamzaoui, Pascal Frossard |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Encoder-Driven Inpainting Strategy in Multiview Video CompressionabstractIn free viewpoint video systems, a user has the freedom to select a virtual view from which an image of the 3D scene is rendered, and the scene is commonly represented by color and depth images of multiple nearby viewpoints. In such representation, there exists data redundancy across multiple dimensions: 1) a 3D voxel may be represented by pixels in multiple viewpoint images (inter-view redundancy); 2) a pixel patch may recur in a distant spatial region of the same image due to self-similarity (inter-patch redundancy); and 3) pixels in a local spatial region tend to be similar (inter-pixel redundancy). It is important to exploit these redundancies during inter-view prediction toward effective multiview video compression. In this paper, we propose an encoder-driven inpainting strategy for inter-view predictive coding, where explicit instructions are transmitted minimally, and the decoder is left to independently recover remaining missing data via inpainting, resulting in lower coding overhead. In particular, after pixels in a reference view are projected to a target view via depth-image-based rendering at the decoder, the remaining holes in the target view are filled via an inpainting process in a block-by-block manner. First, blocks are ordered in terms of difficulty-to-inpaint by the decoder. Then, explicit instructions are only sent for the reconstruction of the most difficult blocks. In particular, the missing pixels are explicitly coded via a graph Fourier transform or a sparsification procedure using discrete cosine transform, leading to low coding cost. For blocks that are easy to inpaint, the decoder independently completes missing pixels via template-based inpainting. We apply our proposed scheme to frames in a prediction structure defined by JCT-3V where inter-view prediction is dominant, and experimentally we show that our scheme achieves up to 3-dB gain in peak-signal-to-noise-ratio in reconstructed image quality over a comparable 3D-High Efficiency Video Coding implementation using fixed 16 $\times $ 16 block size. Yu Gao 0003, Gene Cheung, Thomas Maugey, Pascal Frossard, Jie Liang 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | Reference View Selection in DIBR-Based Multiview CodingabstractAugmented reality, interactive navigation in 3D scenes, multiview video, and other emerging multimedia applications require large sets of images, hence larger data volumes and increased resources compared with traditional video services. The significant increase in the number of images in multiview systems leads to new challenging problems in data representation and data transmission to provide high quality of experience on resource-constrained environments. In order to reduce the size of the data, different multiview video compression strategies have been proposed recently. Most of them use the concept of reference or key views that are used to estimate other images when there is high correlation in the data set. In such coding schemes, the two following questions become fundamental: 1) how many reference views have to be chosen for keeping a good reconstruction quality under coding cost constraints? And 2) where to place these key views in the multiview data set? As these questions are largely overlooked in the literature, we study the reference view selection problem and propose an algorithm for the optimal selection of reference views in multiview coding systems. Based on a novel metric that measures the similarity between the views, we formulate an optimization problem for the positioning of the reference views, such that both the distortion of the view reconstruction and the coding rate cost are minimized. We solve this new problem with a shortest path algorithm that determines both the optimal number of reference views and their positions in the image set. We experimentally validate our solution in a practical multiview distributed coding system and in the standardized 3D-HEVC multiview coding scheme. We show that considering the 3D scene geometry in the reference view, positioning problem brings significant rate-distortion improvements and outperforms the traditional coding strategy that simply selects key frames based on the distance between cameras. Thomas Maugey, Giovanni Petrazzuoli, Pascal Frossard, Marco Cagnazzo, Béatrice Pesquet-Popescu |
IEEE Trans. Image Process. | 1 |
| 2015 | Guided inpainting with cluster-based auxiliary informationabstractIn this paper, we propose a new guided inpainting algorithm based on the exemplar-based approach in order to effectively fill in holes in image synthesis applications. Guided inpainting techniques can be very useful in settings where one has access to the ground truth information like most multiview coding applications. We propose a new auxiliary information based on patch clustering, which is used to refine the candidate exemplar set in the inpainting. For that purpose, a new recursive clustering method based on locally linear embedding (LLE) is introduced. We then design the guided inpainting solution based on LLE with clustered patches, which contrains the reconstruction to operate in one patch cluster only. The index of the appropriate cluster considered as auxiliary information. Experimental results show that our clustering algorithm provides clusters that are well suited to the inpainting problem. They also show that the auxiliary information enables to significantly improve the quality of the inpainted image for a small coding cost. This work is the first study to show that effective inpainting can be performed when the auxiliary information is properly adapted to the characteristics of both the hole and the known texture. Thomas Maugey, Pascal Frossard, Christine Guillemot |
ICIP | 1 |
| 2015 | Universal lossless coding with random user access: The cost of interactivityabstractWe consider the problem of video compression with free viewpoint interactivity. It is well believed that allowing the user to choose its view will incur some loss in terms of compression efficiency. Here we derive the complete rate-storage region for universal lossless coding under the constraint of choosing the view at the receiver. This leads to a counterintuitive result: freely choosing its view at the receiver incurs a loss in terms of storage only and not in the transmission rate. The gain of the optimal scheme with respect to interactive schemes proposed so far is derived and a practical scheme that achieves this gain is proposed. Aline Roumy, Thomas Maugey |
ICIP | 2 |
| 2015 | Optimal layered representation for adaptive interactive multiview video streaming
Ana De Abreu, Laura Toni, Nikolaos Thomos, Thomas Maugey, Fernando Pereira 0001, Pascal Frossard |
J. Vis. Commun. Image Represent. | 4 |
| 2015 | Graph-Based Representation for Multiview Image GeometryabstractIn this paper, we propose a new geometry representation method for multiview image sets. Our approach relies on graphs to describe the multiview geometry information in a compact and controllable way. The links of the graph connect pixels in different images and describe the proximity between pixels in 3D space. These connections are dependent on the geometry of the scene and provide the right amount of information that is necessary for coding and reconstructing multiple views. Our multiview image representation is very compact and adapts the transmitted geometry information as a function of the complexity of the prediction performed at the decoder side. To achieve this, our graph-based representation (GBR) carefully selects the amount of geometry information needed before coding. This is in contrast with depth coding, which directly compresses with losses the original geometry signal, thus making it difficult to quantify the impact of coding errors on geometry-based interpolation. We present the principles of this GBR and we build an efficient coding algorithm to represent it. We compare our GBR approach to classical depth compression methods and compare their respective view synthesis qualities as a function of the compactness of the geometry description. We show that GBR can achieve significant gains in geometry coding rate over depth-based schemes operating at similar quality. Experimental results demonstrate the potential of this new representation. Thomas Maugey, Antonio Ortega, Pascal Frossard |
IEEE Trans. Image Process. | 1 |
| 2015 | Optimized Packet Scheduling in Multiview Video Navigation SystemsabstractWe study coding and transmission strategies in multicamera systems, where correlated sources send data through a bottleneck channel to a central server, which eventually transmits views to different interactive users. We propose a dynamic navigation -path aware packet scheduling optimization under delay, bandwidth, and interactivity constraints aimed at optimizing the quality-of-experience of interactive users. In particular , the scene distortion is minimized jointly with the distortion variations along most likely navigation paths. The optimization relies both on a novel rate-distortion model, which captures the importance of each view in the scene reconstruction , and on an objective function that optimizes resources based on a client navigation model. The latter takes into account the distortion experienced by interactive clients as well as the distortion variations that might be observed by clients during multiview navigation. We solve the scheduling problem with a novel trellis-based solution, which permits to formally decompose the multivariate optimization problem, thereby significantly reducing the computation complexity. Simulation results show the PSNR quality gain offered by the proposed algorithm compared to baseline scheduling policies. Finally, we show that the best scheduling policy consistently adapts to the most likely user navigation path and that it minimizes distortion variations that can be very disturbing for users in traditional navigation systems. Laura Toni, Thomas Maugey, Pascal Frossard |
IEEE Trans. Multim. | 2 |
| 2014 | 3D geometry representation using multiview coding of image tilesabstractCompression of dynamic 3D geometry obtained from depth sensors is challenging, because noise and temporal inconsistency inherent in acquisition of depth data means there is no one-to-one correspondence between sets of 3D points in consecutive time instants. In this paper, instead of coding 3D points (or meshes) directly, we propose to represent an object's 3D geometry as a collection of tile images. Specifically, we first place a set of image tiles around an object. Then, we project the object's 3D geometry onto the tiles that are interpreted as 2D depth images, which we subsequently encode using a modified multiview image codec tuned for piecewise smooth signals. The crux of the tile image framework is the “optimal” placement of image tiles - one that yields the best tradeoff in rate and distortion. We show that if only planar and cylindrical tiles are considered, then the optimal placement problem for K tiles can be mapped to a tractable piece-wise linear approximation problem. We propose an efficient dynamic programming algorithm to find an optimal solution to the piecewise linear approximation problem. Experimental results show that optimal tiling outperforms naïve tiling by up to 35% in rate reduction, and graph transform can further exploit the smoothness of the tile images for coding gain. Yu Gao 0003, Gene Cheung, Thomas Maugey, Pascal Frossard, Jie Liang 0001 |
ICASSP | 3 |
| 2014 | Luminance coding in graph-based representation of multiview imagesabstractMulti-view video transmission poses great challenges because of its data size and dimension. Therefore, how to design efficient 3D scene representations and coding (of luminance and geometry) has become a critical research topic. Recently, the graph-based representation (GBR) is introduced, which provides a lossless compression of multi-view geometry by connecting informative pixels among views. This representation has been shown as a promising alternative to the classical depth-based representation, where the view synthesis accuracy is hard to control. In this work, we study the luminance compression under GBR, which is not well considered in existing literature. With a proper structural reformulation, we show that the graph-based transform can be applied on the GBR paradigm, hence better extracting the correlation among pixels along graph connections. Moreover, we extend the popular SPIHT coding scheme to further improve coding efficiency. The experimental results show that our method leads to better RD coding performance as compared the classical luminance coding algorithms. Thomas Maugey, Yung Hsuan Chao, Akshay Gadde, Antonio Ortega, Pascal Frossard |
ICIP | 1 |
| 2014 | Multiview video representations for quality-scalable navigationabstractInteractive multiview video (IMV) applications offer to users the freedom of selecting their preferred viewpoint. Usually, in these systems texture and depth maps of captured views are available at the user side, as they permit the rendering of intermediate virtual views. However, the virtual views' quality depends on the distance to the available views used as references and on their quality, which is generally constrained by the heterogeneous capabilities of the users. In this context, this work proposes an IMV scalable system, where views are optimally organized in layers, each one offering an incremental improvement in the interactive navigation quality. We propose a distortion model for the rendered virtual views and an algorithm that selects the optimal views' subset per layer. Simulation results show the efficiency of the proposed distortion model, and that the careful choice of reference cameras permits to have a graceful quality degradation for clients with limited capabilities. Ana De Abreu, Laura Toni, Thomas Maugey, Nikolaos Thomos, Pascal Frossard, Fernando Pereira 0001 |
VCIP | 3 |
| 2014 | Key view selection in distributed multiview codingabstractMultiview image and video systems with large number of views lead to new problems in data representation, transmission and user interaction. In order to reduce the data volumes, most distributed multiview coding schemes exploit the inter-view redundancies at the decoder side, using view synthesis from key views. In the situation where many views are considered, the two following questions become fundamental: i) how many key views have to be chosen for keeping a good reconstruction quality with reasonable coding cost? ii) where to place them optimally in the multiview sequences? We propose in this paper an algorithm for selecting the key views in a distributed multiview coding scheme. Based on a novel metric for the correlation between the views, we formulate an optimization problem for the positioning of the key views such that both the distortion of the reconstruction and the coding rate cost are effectively minimized. We then propose a new optimization strategy based on shortest path algorithm that permits to determine both the optimal number of key views and their positions in the image set. We experimentally validate our solution in a practical distributed multiview coding system and we show that considering the 3D scene geometry in the key view positioning brings significant rate-distortion improvements compared to distance-based key view selection as it is commonly done in the literature. Thomas Maugey, Giovanni Petrazzuoli, Pascal Frossard, Marco Cagnazzo, Béatrice Pesquet-Popescu |
VCIP | 1 |
| 2014 | Packet scheduling in multicamera capture systemsabstractIn multiview video services, multiple cameras acquire the same scene from different perspectives, which results in correlated video streams. This generates large amounts of highly redundant data, which need to be properly handled during encoding and transmission of the multi-view data. In this work, we study coding and transmission strategies in multicamera sets, where correlated sources need to be sent to a central server through a bottleneck channel, and eventually delivered to interactive clients. We propose a dynamic correlation-aware packet scheduling optimization under delay, bandwidth, and interactivity constraints. A novel trellis-based solution permits to formally decompose the multivariate optimization problem, thereby significantly reducing the computation complexity. Simulation results show the gain of the proposed algorithm compared to baseline scheduling policies. Laura Toni, Thomas Maugey, Pascal Frossard |
VCIP | 2 |
| 2014 | Extended Layered Depth Image Representation in Multiview NavigationabstractEmerging applications in multiview streaming look for providing interactive navigation services to video players. The user can ask for information from any viewpoint with a minimum transmission delay. The purpose is to provide user with as much information as possible with least number of redundancies. The recent concept of navigation segment representation consists of regrouping a given number of viewpoints in one signal and transmitting them to the users according to their navigation path. The question of the best description strategy of these navigation segments is however still open. In this paper, we propose to represent and code navigation segments by a method that extends the recent layered depth image (LDI) format. It consists of describing the scene from a viewpoint with multiple images organized in layers corresponding to the different levels of occluded objects. The notion of extended LDI comes from the fact that the size of this image is adapted to take into account the sides of the scene also, in contrary to classical LDI. The obtained results show a significant rate-distortion gain compared to classical multiview compression approaches in navigation scenario. Uday Takyar, Thomas Maugey, Pascal Frossard |
IEEE Signal Process. Lett. | 2 |
| 2014 | Depth-Based Multiview Distributed Video CodingabstractMultiview distributed video coding (DVC) has gained much attention in the last few years because of its potential in avoiding communication between cameras without decreasing the coding performance. However, the current results are not matching the expectations mainly due to the fact that some theoretical assumptions are not satisfied in the current implementations. For example, in distributed source coding the encoder must know the correlation between the sources, which cannot be achieved in the traditional DVC systems without having a communication between the cameras. In this work, we propose a novel multiview distributed video coding scheme in which the depth maps are used to estimate the way two views are correlated with no exchanges between the cameras. Only their relative positions are known. We design the complete scheme and further propose a rate allocation algorithm to efficiently share the bit budget between the different components of our scheme. Then, a rate allocation algorithm for depth maps is proposed in order to maximize the quality of synthesized virtual views. We show, through detailed experiments, that our scheme significantly outperforms the state-of-the-art DVC system. Giovanni Petrazzuoli, Thomas Maugey, Marco Cagnazzo, Béatrice Pesquet-Popescu |
IEEE Trans. Multim. | 2 |
| 2014 | Correlation-Aware Packet Scheduling in Multi-Camera NetworksabstractIn multiview applications, multiple cameras acquire the same scene from different viewpoints and generally produce correlated video streams. This results in large amounts of highly redundant data. In order to save resources, it is critical to handle properly this correlation during encoding and transmission of the multiview data. In this work, we propose a correlation-aware packet scheduling algorithm for multi-camera networks, where information from all cameras are transmitted over a bottleneck channel to clients that reconstruct the multiview images. The scheduling algorithm relies on a new rate-distortion model that captures the importance of each view in the scene reconstruction. We propose a problem formulation for the optimization of the packet scheduling policies, which adapt to variations in the scene content. Then, we design a low complexity scheduling algorithm based on a trellis search that selects the subset of candidate packets to be transmitted towards effective multiview reconstruction at clients. Extensive simulation results confirm the gain of our scheduling algorithm when inter-source correlation information is used in the scheduler, compared to scheduling policies with no information about the correlation or non-adaptive scheduling policies. We finally show that increasing the optimization horizon in the packet scheduling algorithm improves the transmission performance, especially in scenarios where the level of correlation rapidly varies with time. Laura Toni, Thomas Maugey, Pascal Frossard |
IEEE Trans. Multim. | 2 |
| 2013 | Graph-based representation and coding of multiview geometryabstractWe propose a new approach for describing the geometry information of multiview image representations. Rather than transmitting the raw geometry of the scene, under the form of depth information, we build a graph that represents the connections between corresponding pixels in different views in a multiview image set. The graph starts with the reference image and recursively represents the next levels (i.e., images) by storing the new pixels (those that cannot be derived from the previous image) and their connections to the lower level. The decoder uses these connections to recover the multiple images. In addition to being natural and more easily controlled, the proposed graph-based representation can be compressed more efficiently than depth images. This new representation offers promising perspectives for effective and flexible coding in multiview imaging. Thomas Maugey, Antonio Ortega, Pascal Frossard |
ICASSP | 1 |
| 2013 | Bayesian Early Mode Decision Technique for View Synthesis Prediction-Enhanced Multiview Video CodingabstractView synthesis prediction (VSP) is a coding mode that predicts video blocks from synthesised frames. It is particularly useful in a multi-camera setup with large inter-camera distances. Adding a VSP-based SKIP mode to a standard Multiview Video Coding (MVC) framework improves the rate-distortion (RD) performance but increases the time complexity of the encoder. This letter proposes an early mode decision technique for VSP SKIP-enhanced MVC. Our method uses the correlation between the RD costs of the VSP SKIP mode in neighbouring views and Bayesian decision theory to reduce the number of candidate coding modes for a given macroblock. Simulation results showed that our technique can save up to 36.20% of the encoding time without any significant loss in RD performance. Shadan Khan Khattak, Raouf Hamzaoui, Thomas Maugey, Pascal Frossard |
IEEE Signal Process. Lett. | 3 |
| 2013 | Evaluation of Side Information Effectiveness in Distributed Video CodingabstractThe rate-distortion performance of a distributed video coding system strongly depends on the characteristics of the side information. One could naïvely think that the best side information is the one with the largest PSNR with respect to the original corresponding image. However, previous works have shown that this is not always the case and a reduction of the side information MSE does not always translate into better rate-distortion performance for the complete system. The scope of this paper is to explore a set of metrics other than the PSNR and explicitly designed to classify the side information with respect to its impact on the end-to-end compression performance. A first contribution is to define an experimental framework that can be used to meaningfully compare different metrics for side information evaluation. As a second contribution, our analysis allows to understand why in some cases PSNR-based metrics provide a fairly reliable estimation of the side information quality, while in other cases they do not. This analysis also allows us to introduce a set of new metrics that are better adapted for side information effectiveness evaluation, and that are based on a suitable power of the absolute difference between side information and the original image, or on the Hamming distance between the respective transform coefficients. Besides their theoretical interest, these new metrics can also improve the rate-distortion performance of some distributed video coding systems such as the hash-based ones. We observe improvement up to 74% rate reduction in a simple study case. Thomas Maugey, Jérôme Gauthier, Marco Cagnazzo, Béatrice Pesquet-Popescu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Navigation Domain Representation For Interactive Multiview ImagingabstractEnabling users to interactively navigate through different viewpoints of a static scene is a new interesting functionality in 3D streaming systems. While it opens exciting perspectives toward rich multimedia applications, it requires the design of novel representations and coding techniques to solve the new challenges imposed by the interactive navigation. In particular, the encoder must prepare a priori a compressed media stream that is flexible enough to enable the free selection of multiview navigation paths by different streaming media clients. Interactivity clearly brings new design constraints: the encoder is unaware of the exact decoding process, while the decoder has to reconstruct information from incomplete subsets of data since the server generally cannot transmit images for all possible viewpoints due to resource constrains. In this paper, we propose a novel multiview data representation that permits us to satisfy bandwidth and storage constraints in an interactive multiview streaming system. In particular, we partition the multiview navigation domain into segments, each of which is described by a reference image (color and depth data) and some auxiliary information. The auxiliary information enables the client to recreate any viewpoint in the navigation segment via view synthesis. The decoder is then able to navigate freely in the segment without further data request to the server; it requests additional data only when it moves to a different segment. We discuss the benefits of this novel representation in interactive navigation systems and further propose a method to optimize the partitioning of the navigation domain into independent segments, under bandwidth and storage constraints. Experimental results confirm the potential of the proposed representation; namely, our system leads to similar compression performance as classical inter-view coding, while it provides the high level of flexibility that is required for interactive streaming. Because of these unique properties, our new framework represents a promising solution for 3D data representation in novel interactive multimedia services. Thomas Maugey, Ismaël Daribo, Gene Cheung, Pascal Frossard |
IEEE Trans. Image Process. | 1 |
| 2013 | Interactive Multiview Video System With Low Complexity 2D Look Around at DecoderabstractMultiview video with interactive 2D look around at the receiver is a challenging application with several issues in terms of effective use of storage and bandwidth resources, reactivity of the system, quality of the viewing experience and system complexity. The impression of 3D immersion is highly dependent on the smoothness of the navigation and thus on the number of 2D viewpoints. The classical decoding system for generating virtual views first projects a reference or encoded frame to a given viewpoint and then fills in the holes due to potential occlusions. This last step still constitutes a complex operation with specific software or hardware at the receiver and requires a certain quantity of information from the neighboring frames for ensuring consistency between the virtual images. In this work we propose a new approach that shifts most of the burden due to interactivity from the decoder to the encoder, by anticipating the navigation of the decoder and sending auxiliary information that guarantees temporal and interview consistency. This leads to an additional cost in terms of transmission rate and storage, which we minimize by using optimization techniques based on the user behavior modeling. We show by experiments that the proposed system represents a valid solution for interactive multiview systems with classical decoders. Thomas Maugey, Pascal Frossard |
IEEE Trans. Multim. | 1 |
| 2012 | Consistent view synthesis in interactive multiview imagingabstractAn important question in the design of interactive multiview systems consists in determining the information needed by the decoder for high quality navigation between the views. Most of the existing techniques focus on the captured sequences and only consider their transmission, which does not guarantee consistency among receiver-generated frames of chosen virtual views. In this work, we propose a solution that additional transmits auxiliary information in order to help the construction of synthesized views, especially in the occluded areas. Comparative results with existing approaches validate this novel representation of multiview data for interactive navigation. We show that decoding quality and consistency among frames are improved with only a small share of additional information. Thomas Maugey, Pascal Frossard, Gene Cheung |
ICIP | 1 |
| 2011 | Rate distorsion analysis in a disparity compensated schemeabstractThis paper addresses the problem of rate distortion analysis in the context of multi-view image coding, where images are predicted via disparity compensation based on depth map. We first present an analytical model for the variance of the residual error in a predicted frame when the prediction is done with the help of a compressed depth map. This residual variance model presents a convenient expression that separates the different error origins (reference frame quantization, depth map coding, and motion activity). We then validate the novel analytical model by testing separately its different underlying hypotheses. Finally, we illustrate an application of our analytical model in a simple bit allocation problem where the objective is to determine the optimal distribution of a global bit budget among reference frame, depth map and disparity-compensated frame. We observe that the optimal allocation given by the analytical model corresponds in practice to the best rate distribution for high bitrate, which confirms the potential of the proposed model in the design of rate-controlled multi-view coding algorithms. Valentina Davidoiu, Thomas Maugey, Béatrice Pesquet-Popescu, Pascal Frossard |
ICASSP | 2 |
| 2011 | Interactive multiview video system with low decoding complexityabstractResearch in multimedia is always investigating new ways of improving the immersive experience of the users. One current solution consists in designing systems which offer a high level of interactivity, such as multiview content navigation where the point of view can be changed while watching at a video sequence (e.g., free view- point television, gaming, etc.). The coding algorithm designed for the transmission of such media streams must be adapted to these novel decoder needs. However, video plus depth data transmission is usually performed by considering the information flows as two sequences encoded with MVC schemes. Whereas it achieves good compression performance, this coding approach is not appropriate for interactive applications since the decoding of a frame of- ten requires the prior transmission and decoding of several reference frames. Moreover, the techniques recently developed to improve interactivity are generally implemented at the decoder, whose computational complexity requirements are augmented. In this paper, we propose a novel coding scheme for video plus depth sequences that is adapted to user navigation; contrarily to several common approaches, the additional complexity is added on the encoder side so that the decoder stays simple. We further propose to limit the additional bandwidth imposed by interactivity requirements by designing a rate allocation algorithm that builds on a model of the user behavior. A first version of our novel coding architecture is evaluated in terms of rate-distortion performance, where it is shown to offer a high interactivity at a reasonable bandwidth cost. Thomas Maugey, Pascal Frossard |
ICIP | 1 |
| 2010 | Using an exponential power model forwyner ziv video codingabstractThe Laplacian model is the standard distribution for correlation noise estimation at the turbodecoder in Wyner-Ziv coding schemes. In practice, this hypothesis is not always satisfied and, regularly, the estimated model sensibly differs from the error distribution. In this work, we prove that using a model better fitted to the true distribution improves the performances, and we thus propose to use the more general exponential power distribution (EPD) which has never been tested in a distributed video coding context. Gains in rate-distortion over the Laplacian model are illustrated by results on several video sequences, showing that the EPD model outperforms the Laplacian one in off-line (oracle) as well as in on-line (practical implementation) modes. These results also indicate that, in some cases, the online EPD model reduces the bitrate even over the off-line Laplacian model. Thomas Maugey, Jérôme Gauthier, Béatrice Pesquet-Popescu, Christine Guillemot |
ICASSP | 1 |
| 2010 | Compressed sensing of multiview images using disparity compensationabstractCompressed sensing is applied to multiview image sets and inter-image disparity compensation is incorporated into image reconstruction in order to take advantage of the high degree of inter-image correlation common to multiview scenarios. Instead of recovering images in the set independently from one another, two neighboring images are used to calculate a prediction of a target image, and the difference between the original measurements and the compressed-sensing projection of the prediction is then reconstructed as a residual and added back to the prediction in an iterated fashion. The proposed method shows large gains in performance over straightforward, independent compressed-sensing recovery. Additionally, projection and recovery are block-based to significantly reduce computation time. Maria Trocan, Thomas Maugey, Eric W. Tramel, James E. Fowler, Béatrice Pesquet-Popescu |
ICIP | 2 |
| 2010 | Disparity-compensated compressed-sensing reconstruction for multiview imagesabstractIn a multiview-imaging setting, image-acquisition costs could be substantially diminished if some of the cameras operate at a reduced quality. Compressed sensing is proposed to effectuate such a reduction in image quality wherein certain images are acquired with random measurements at a reduced sampling rate via projection onto a random basis of lower dimension. To recover such projected images, compressed-sensing recovery incorporating disparity compensation is employed. Based on a recent compressed-sensing recovery algorithm for images that couples an iterative projection-based reconstruction with a smoothing step, the proposed algorithm drives image recovery using the projection-domain residual between the random measurements of the image in question and a disparity-based prediction created from adjacent, high-quality images. Experimental results reveal that the disparity-based reconstruction significantly outperforms direct reconstruction using simply the random measurements of the image alone. Maria Trocan, Thomas Maugey, James E. Fowler, Béatrice Pesquet-Popescu |
ICME | 2 |
| 2010 | Side information enhancement using an adaptive hash-based genetic algorithm in a Wyner-Ziv contextabstractSide information construction in Wyner-Ziv video coding is a sensible task which strongly influences the final ratedistortion performance of the scheme. This side information is usually generated through an interpolation of the previous and next images. Some of the zones of a scene however, such as the occlusions, cannot be estimated with other frames. In this paper we propose to avoid this problem by sending some hash information for these unpredictable zones of the image. The resulting algorithm is described and tested here. The obtained results show the advantages of using localized hash information for the high error zones in distributed video coding. Thomas Maugey, Charles Yaacoub, Joumana Farah, Marco Cagnazzo, Béatrice Pesquet-Popescu |
MMSP | 1 |
| 2010 | Side information refinement for long duration GOPs in DVCabstractSide information generation is a critical step in distributed video coding systems. This is performed by using motion compensated temporal interpolation between two or more key frames (KFs). However, when the temporal distance between key frames increases (i.e. when the GOP size becomes large), the linear interpolation becomes less effective. In a previous work we showed that this problem can be mitigated by using high order interpolation. Now, in the case of long duration GOP, state-of-the-art algorithms propose a hierarchical algorithm for side information generation. By using this procedure, the quality of the central interpolated image in a GOP is consistently worse than images closer to the KFs. In this paper we propose a refinement of the central WZFs by higher order interpolation of the already decoded WZFs, that are closer to the WZF to be estimated. So we reduce the fluctuation of side information quality, with a beneficial impact on final rate-distortion characteristics of the system. The experimental results show an improvement on the SI up to 2.71 dB with respect the state-of-the-art and a global improvement of the PSNR on the decoded frames up to 0.71 dB and a bit rate reduction up to 15%. Giovanni Petrazzuoli, Thomas Maugey, Marco Cagnazzo, Béatrice Pesquet-Popescu |
MMSP | 2 |
| 2010 | Multistage compressed-sensing reconstruction of multiview imagesabstractCompressed sensing is applied to multiview image sets and the high degree of correlation between views is exploited to enhance recovery performance over straightforward independent view recovery. This gain in performance is obtained by recovering the difference between a set of acquired measurements and the projection of a prediction of the signal they represent. The recovered difference is then added back to the prediction, and the prediction and recovery procedure is repeated in an iterated fashion for each of the views in the multiview image set. The recovered multiview image set is then used as an initialization to repeat the entire process again to form a multistage refinement. Experimental results reveal substantial performance gains from the multistage reconstruction. Maria Trocan, Thomas Maugey, Eric W. Tramel, James E. Fowler, Béatrice Pesquet-Popescu |
MMSP | 2 |
| 2009 | A differential motion estimation method for image interpolation in distributed video codingabstractMotion estimation methods based on differential techniques proved to be very useful in the context of video analysis, but have a limited employment in classical video compression because, though accurate, the dense motion vector field they produce requires too much coding resource and computational effort. On the contrary, this kind of algorithm could be useful in the framework of distributed video coding (DVC). In this paper we propose a differential motion estimation algorithm which can run at the decoder in a DVC scheme, without requiring any increase in coding rate. This algorithm allows a performance improvement in image interpolation with respect to state-of-the-art algorithms. Marco Cagnazzo, Thomas Maugey, Béatrice Pesquet-Popescu |
ICASSP | 2 |
| 2009 | Dense disparity estimation in a multi-view distributed video coding systemabstractDistributed video coding (DVC) is a recent paradigm which aims at transferring part of the coding complexity from the encoder to the decoder. The performance of such a coding scheme strongly depends on the capacity to estimate correlation at the decoder and, consequently, on the side information quality. In this paper we consider a multi-view DVC framework and propose a very efficient dense disparity estimation technique for side information construction, based on a variational formulation. The simulation results show that our approach clearly outperforms the existing methods for inter-view side-information generation. Thomas Maugey, Wided Miled, Béatrice Pesquet-Popescu |
ICASSP | 1 |
| 2009 | Image interpolation with edge-preserving differential motion refinementabstractMotion estimation (ME) methods based on differential techniques provide useful information for video analysis, and moreover it is relatively easy to embed into them regularity constraints enforcing for example, contour preservation. On the other hand, these techniques are rarely employed for video compression since, though accurate, the dense motion vector field (MVF) they produce requires too much coding resource and computational effort. However, this kind of algorithm could be useful in the framework of distributed video coding (DVC), where the motion vector are computed at the decoder side, so that no bit-rate is needed to transmit them. Moreover usually the decoder has enough computational power to face with the increased complexity of differential ME. In this paper we introduce a new image interpolation algorithm to be used in the context of DVC. This algorithm combines a popular DVC technique with differential ME. We adapt a pel-recursive differential ME algorithm to the DVC context; moreover we insert a regularity constraint which allows more consistent MVFs. The experimental results are encouraging: the quality of interpolated images is improved of up to 1.1 dB w.r.t. to state-of-the-art techniques. These results prove to be consistent when we use different GOP sizes. Marco Cagnazzo, Wided Miled, Thomas Maugey, Béatrice Pesquet-Popescu |
ICIP | 3 |
| 2008 | Side information estimation and new symmetric schemes for multi-view distributed video coding
Thomas Maugey, Béatrice Pesquet-Popescu |
J. Vis. Commun. Image Represent. | 1 |