EDBT 2026 Demo / reviewers in the wild / expert
Christine Guillemot
dblp:40/1951 · also Christine M. Guillemot
· DBLP profile ↗
232ranked-venue papers
10as first author
36since 2021 · last 2026
0000-0003-1604-967XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 203 · 8 first-author · 31 since 2021Computer networks · 12 · 1 first-authorArtificial intelligence and machine learning · 11 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Databases, data management, data science and information retrieval · 5Theory of computation · 3Systems, architecture and hardware · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAFe-GS: Compactness-Aware Frequency-Guided Densification for 3D Gaussian Splatting
Léo-Paul Huar, Gustavo L. Sandri, Neus Sabater, Christine Guillemot, Pierre Hellier |
ICPR (4) | 4 |
| 2026 | DUALF-D: Disentangled dual-hyperprior approach for light field image compressionabstractLight field (LF) imaging captures spatial and angular information, offering a 4D scene representation enabling enhanced visual understanding. However, high dimensionality and redundancy across spatial and angular domains present major challenges for compression, particularly where storage, transmission bandwidth, or processing latency are constrained. We present a novel Variational Autoencoder (VAE)-based framework that explicitly disentangles spatial and angular features using two parallel latent branches. Each branch is coupled with an independent hyperprior model, allowing more precise distribution estimation for entropy coding and finer rate–distortion control. This dual-hyperprior structure enables the network to adaptively compress spatial and angular information based on their unique statistical characteristics, improving coding efficiency. To further enhance latent feature specialization and promote disentanglement, we introduce a mutual information-based regularization term that minimizes redundancy between the two branches while preserving feature diversity. Unlike prior methods relying on covariance-based penalties prone to collapse, our information-theoretic regularizer provides more stable and interpretable latent separation. Experimental results on publicly available LF datasets demonstrate our method achieves strong compression performance, yielding an average BD-PSNR gain of 2.91 dB over HEVC and high compression ratios (e.g., 200:1). Additionally, our design enables fast inference, with a total end-to-end time over 19x faster than the JPEG Pleno standard, making it well-suited for real-time and bandwidth-sensitive applications. By jointly leveraging disentangled representation learning, dual-hyperprior modeling, and information-theoretic regularization, our approach offers a scalable, effective solution for practical light field image compression. Soheib Takhtardeshir, Roger Olsson, Christine Guillemot, Mårten Sjöström |
Signal Process. Image Commun. | 3 |
| 2026 | Visibility-Based Geometry Pruning of Neural Plenoptic Scene RepresentationsabstractThe need for more realistic 3D scene representations has fomented the development of models for a wide range of applications. In this context, solutions that attempt to model the light's behavior through the plenoptic function have provided considerable advancements using neural-based approaches, often presenting a trade-off between rendering time and model sizes. In this work, we propose a pruning framework to reduce the sizes of these models by computing the visibility over the training data, applicable to different 3D scene representations. In particular, we implement first a solution suitable for the 3D Gaussian Splatting, and then we exemplify the solution for the Neural Radiance Fields (NeRF)-style of rendering using PlenOctrees. We show that our pruning solution produces smaller models in terms of the number of elements – be they voxels, points, or Gaussians – with minimal losses in terms of rendering novel views. We further assess our solution by combining it with state-of-the-art (SOTA) compression solutions for both rendering schemes. Results over the NeRF-Synthetic dataset show comparable metrics to the SOTA for PlenOctrees, achieving marginal gains for lower bitrates. For 3DGS, the combination of our pruning method and compression solutions achieves a compression ratio of up to 37.5 times over the uncompressed 3DGS models, with only a 0.5 dB decrease in rendering quality. When compared against other SOTA compression methods, our solution produces models 1.4 times smaller, with less than a 0.1 dB loss over novel views for synthetic data, and models 1.9 times smaller with less than 0.2 dB loss when synthesizing novel views on real-world, outdoor content. Davi Rabbouni Freitas, Ioan Tabus, Christine Guillemot |
IEEE Trans. Multim. | 3 |
| 2025 | DUALF-C: Disentangled Light Field Compression with Entropy-Aware Bitstream GenerationabstractRecent advancements in learned light field (LF) image compression highlight the advantages of modeling spatial and angular redundancies using deep generative models. Among these, Variational Autoencoders (VAEs) have shown strong potential in learning compact latent representations of LF data. However, many existing approaches rely on simple uniform quantization and basic entropy coding, which limits compression efficiency in practical applications. This paper introduces DUALF-C, a lightweight and retraining-free compression pipeline that augments pretrained VAE-based LF compression models using structured post-encoding transformations. The proposed framework integrates bitplane slicing, latent channel reordering, non-uniform quantization, and patch-based vector quantization to improve bitrate efficiency while preserving reconstruction quality. Experimental evaluations demonstrate that DUALF-C significantly reduces bit-per-pixel (BPP) without degrading image quality, making it a practical solution for bandwidth-constrained immersive imaging systems. Soheib Takhtardeshir, Roger Olsson, Christine Guillemot, Mårten Sjöström |
VCIP | 3 |
| 2024 | A Comparative Assessment of Implicit and Explicit Plenoptic Scene Representationsabstract3D scene representation has been a central theme of study for a wide range of applications, and the representation of light behavior is one of the relevant topics when producing realistic models. In this work, we create a framework to assess the representation of non-Lambertian scenes by generating a pipeline to create plenoptic point clouds (PPCs) systematically and evaluating them against implicit solutions, such as Neural Radiance Fields (NeRF)-like models. We compare such approaches according to rendering quality and compression efficiency. On the compression side, we propose an encoding scheme for PPC, leveraging the occlusion masks of the points and the Moving Picture Expert Group's (MPEG) Geometry-Based Solid Content Test Model (GeS-TM). Rendering results over the training views show that the uncompressed PPC outperforms 3D Gaussian Splatting (3DGS) by 1.51 dB, on average, for the 8 scenes of the NeRF Synthetic 360 dataset. In compression efficiency, 3DGS outperforms the compressed PPCs by 0.7 dB in BD-PSNR on average. Our occlusion-aware encoding scheme reduces the size of uncompressed PPCs up to 800 times, outperforming current encoding schemes for PPC by 1.9 dB in BD-PSNR. Davi Rabbouni Freitas, Ricardo L. de Queiroz, Ioan Tabus, Christine Guillemot |
MMSP | 4 |
| 2024 | Learning Kernel-Modulated Neural Representation for Efficient Light Field CompressionabstractLight fields capture 3D scene information by recording light rays emitted from a scene at various orientations. They offer a more immersive perception, compared with classic 2D images, but at the cost of huge data volumes. In this paper, we design a compact neural network representation for the light field compression task. In the same vein as the deep image prior, the neural network takes randomly initialized noise as input and is trained in a supervised manner in order to best reconstruct the target light field Sub-Aperture Images (SAIs). The network is composed of two types of complementary kernels: descriptive kernels (descriptors) that store scene description information learned during training, and modulatory kernels (modulators) that control the rendering of different SAIs from the queried perspectives. To further enhance compactness of the network meanwhile retain high quality of the decoded light field, we propose modulator allocation and apply kernel tensor decomposition techniques, followed by non-uniform quantization and lossless entropy coding. Extensive experiments demonstrate that our method outperforms other state-of-the-art (SOTA) methods by a significant margin in the light field compression task. Moreover, after adapting descriptors, the modulators learned from one light field can be transferred to new light fields for rendering dense views, showing the potential of the solution for view synthesis. Jinglei Shi, Christine Guillemot |
IEEE Trans. Image Process. | 3 |
| 2023 | Joint Compression and Demosaicking For Satellite ImagesabstractImage sensors used in real camera systems are equipped with colour filter arrays which sample the light rays in different spectral bands. Each colour channel can thus be obtained sep-arately by considering the corresponding colour filter. While existing compression solutions mostly assume that the captured raw data has been demosaicked prior to compression, in this paper, we describe an end-to-end trainable neural network for joint compression and demosaicking of satellite images. We first introduce a training loss combining a perceptual loss with the classical mean square error, which is shown to better preserve the high-frequency details present in satellite images. We then present a multi-loss balancing strategy which significantly improves the performance of the proposed joint demosaicking-compression solution. Pascal Bacchus, Renaud Fraisse, Aline Roumy, Christine Guillemot |
ICASSP | 4 |
| 2023 | Unrolled Fourier Disparity Layer Optimization for Scene Reconstruction from Few-Shots Focal StacksabstractThis paper presents a novel unrolled optimization method to reconstruct a dense light field from a focal stack containing only very few images captured with different focus. The proposed unrolled method first reconstructs Fourier Disparity Layers (FDL) from which all the light field viewpoints can then be computed. By recovering details in regions that are out-of-focus in all the captured images, the produced FDL model is also suitable for post-capture scene refocusing from a sparse focal stack. Solving the optimization problem in the FDL domain allows us to derive a closed-form expression of the data-fit term of the inverse problem. We show that the proposed framework outperforms state-of-the-art methods from focal stack measurements for both light field reconstruction and image refocusing. Brandon Le Bon, Mikael Le Pendu, Christine Guillemot |
ICASSP | 3 |
| 2023 | Joint Neural Representation for Multiple Light FieldsabstractNeural implicit representations have appeared as a promising technique for representing a variety of signals, among which light fields. These representations offer several advantages over traditional grid-based representations, such as independence to the signal resolution. Some work has been done to find good initial representations for a given type of signal, usually via meta-learning approaches. However, exploiting the features shared between different scenes remains an understudied problem. We provide a step towards this end by presenting a method for sharing the representation between thousands of light fields, splitting the representation between a part that is shared between all light fields and a part which varies individually from one light field to another. We show that this joint representation possesses good interpolation properties, and allows for a more light-weight storage of a whole database of light fields, exhibiting a ten-fold reduction in the size of the representation when compared to using a separate representation for each light field. Guillaume Le Guludec, Christine Guillemot |
ICASSP | 2 |
| 2023 | Light Field Compression Via Compact Neural Scene RepresentationabstractIn this paper, we propose a novel light field compression method based on a low rank-constrained neural scene representation. While most existing methods directly compress the light field views, our method first learns a Multi-Layer Perceptron (MLP)-based Neural Radiance Field (NeRF) from the input views. To be able to efficiently compress the NeRF scene representation, the weights of the MLP are optimized under a low-rank constraint using the Alternating Direction Method of Multipliers (ADMM) optimization method. The weights of NeRF are then decomposed into Tensor Train (TT) components which allow us to distill original NeRF network into a slimmer one. The slim NeRF is then refined using a quantization-aware training procedure. Experimental results show that this low rank-constrained NeRF-based light field compression method can achieve better rate-distortion than reference methods, while keeping the free-viewpoint reconstruction capability. Jinglei Shi, Christine Guillemot |
ICASSP | 2 |
| 2023 | Filtered Residual Compression for Satellite ImagesabstractLearned image compression neural networks have difficulties adapting to certain satellite image characteristics, especially high frequencies that disappear at a high bit-rate in the blur generated in the reconstruction. To answer this problem we describe a joint end-to-end trainable neural network. It is separated into a general compression network and a smaller specialised network. We train a specialized network to compress the residual part of the image to best preserve the high-frequency details present in the satellite images. The proposed model achieves higher rate-distortion performance than current lossy image compression standards and also manages to retrieve details previously poorly reconstructed. Pascal Bacchus, Renaud Fraisse, Christine Guillemot, Aline Roumy |
IGARSS | 3 |
| 2023 | Vanishing Point Aided Hash-Frequency Encoding for Neural Radiance Fields (NeRF) from Sparse 360°InputabstractNeural Radiance Fields (NeRF) enable novel view synthesis of 3D scenes when trained with a set of 2D images. One of the key components of NeRF is the input encoding, i.e. mapping the coordinates to higher dimensions to learn high-frequency details, which has been proven to increase the quality. Among various input mappings, hash encoding is gaining increasing attention for its efficiency. However, its performance on sparse inputs is limited. To address this limitation, we propose a new input encoding scheme that improves hash-based NeRF for sparse inputs, i.e. few and distant cameras, specifically for 360° view synthesis. In this paper, we combine frequency encoding and hash encoding and show that this combination can increase dramatically the quality of hash-based NeRF for sparse inputs. Additionally, we explore scene geometry by estimating vanishing points in omnidirectional images (ODI) of indoor and city scenes in order to align frequency encoding with scene structures. We demonstrate that our vanishing point-aided scene alignment further improves deterministic and non-deterministic encodings on image regression and NeRF tasks where sharper textures and more accurate geometry of scene structures can be reconstructed. Thomas Maugey, Sebastian Knorr, Christine Guillemot |
ISMAR | 4 |
| 2023 | ZEPI-Net: Light Field Super Resolution via Internal Cross-Scale Epipolar Plane Image Zero-Shot Learning
Zhaolin Xiao, Yinhai Liu, Haiyan Jin, Christine Guillemot |
Neural Process. Lett. | 4 |
| 2023 | PnP-ReG: Learned Regularizing Gradient for Plug-and-Play Gradient DescentabstractAbstract. The plug-and-play framework makes it possible to integrate advanced image denoising priors into optimization algorithms to efficiently solve a variety of image restoration tasks generally formulated as maximum a posteriori (MAP) estimation problems. The plug-and-play alternating direction method of multipliers (ADMM) and the regularization by denoising (RED) algorithms are two examples of such methods that made a breakthrough in image restoration. However, the former plug-and-play approach only applies to proximal algorithms. And while the explicit regularization in RED can be used in various algorithms, including gradient descent, the gradient of the regularizer computed as a denoising residual leads to several approximations of the underlying image prior in the MAP interpretation of the denoiser. We show that it is possible to train a network directly modeling the gradient of a MAP regularizer while jointly training the corresponding MAP denoiser. We use this network in gradient-based optimization methods and obtain better results compared to other generic plug-and-play approaches. We also show that the regularizer can be used as a pretrained network for unrolled gradient descent. Lastly, we show that the resulting denoiser allows for a better convergence of the plug-and-play ADMM. Rita Fermanian, Mikael Le Pendu, Christine Guillemot |
SIAM J. Imaging Sci. | 3 |
| 2023 | Preconditioned Plug-and-Play ADMM with Locally Adjustable Denoiser for Image RestorationabstractAbstract. Plug-and-Play priors recently emerged as a powerful technique for solving inverse problems by plugging a denoiser into a classical optimization algorithm. The denoiser accounts for the regularization and therefore implicitly determines the prior knowledge on the data, hence replacing typical handcrafted priors. In this paper, we extend the concept of Plug-and-Play priors to use denoisers that can be parameterized for nonconstant noise variance. In that aim, we introduce a preconditioning of the ADMM algorithm, which mathematically justifies the use of such an adjustable denoiser. We additionally propose a procedure for training a convolutional neural network for high quality nonblind image denoising that also allows for pixelwise control of the noise standard deviation. We show that our pixelwise adjustable denoiser, along with a suitable preconditioning strategy, can further improve the Plug-and-Play ADMM approach for several applications, including image completion, interpolation, demosaicing, and Poisson denoising. Mikael Le Pendu, Christine Guillemot |
SIAM J. Imaging Sci. | 2 |
| 2023 | Compression of Plenoptic Point Cloud Attributes Using 6-D Point Clouds and 6-D TransformsabstractIn this paper, we introduce a novel 6-D representation of plenoptic point clouds, enabling joint, non-separable transform coding of plenoptic signals defined along both spatial and angular (viewpoint) dimensions. This 6-D representation, which is built in a global coordinate system, can be used in both multi-camera studio capture and video fly-by capture scenarios, with various viewpoint (camera) arrangements and densities. We show that both the Region-Adaptive Hierarchical Transform (RAHT) and the Graph Fourier Transform (GFT) can be extended to the proposed 6-D representation to enable the non-separable transform coding. Our method is applicable to plenoptic data with either dense or sparse sets of viewpoints, and tocompleteorincompleteplenoptic data, while the state-of-the-art RAHT-KLT method, which is separable in spatial and angular dimensions, is applicable only tocompleteplenoptic data. The “complete” plenoptic data refers to data that has, for each spatial point, one colour for every viewpoint (ignoring any occlusions), while “incomplete” data has colours only for thevisiblesurface points at each viewpoint. We demonstrate that the proposed 6-D RAHT and 6-D GFT compression methods are able to outperform the state-of-the-art RAHT-KLT method on 3-D objects with various levels of surface specularity, and captured with different camera arrangements and different degrees of viewpoint sparsity. Maja Krivokuca, Ehsan Miandji, Christine Guillemot, Philip A. Chou |
IEEE Trans. Multim. | 3 |
| 2023 | HEADSET: Human Emotion Awareness under Partial Occlusions Multimodal DataSETabstractThe volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications. Particularly in the context of the rapid advancement of Extended Reality (XR) applications, this volumetric data has proven to be an essential technology for future XR elaboration. In this work, we present a new multimodal database to help advance the development of immersive technologies. Our proposed database provides ethically compliant and diverse volumetric data, in particular 27 participants displaying posed facial expressions and subtle body movements while speaking, plus 11 participants wearing head-mounted displays (HMDs). The recording system consists of a volumetric capture (VoCap) studio, including 31 synchronized modules with 62 RGB cameras and 31 depth cameras. In addition to textured meshes, point clouds, and multi-view RGB-D data, we use one Lytro Illum camera for providing light field (LF) data simultaneously. Finally, we also provide an evaluation of our dataset employment with regard to the tasks of facial expression classification, HMDs removal, and point cloud reconstruction. The dataset can be helpful in the evaluation and performance testing of various XR algorithms, including but not limited to facial expression recognition and reconstruction, facial reenactment, and volumetric video. HEADSET and its all associated raw data and license agreement will be publicly available for research purposes. Fatemeh Ghorbani Lohesara, Davi Rabbouni Freitas, Christine Guillemot, Karen Egiazarian, Sebastian Knorr |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | A New Regularization for Retinex Decomposition of Low-Light ImagesabstractWe study unsupervised Retinex decomposition for low light image enhancement. Being an underdetermined problem with infinite solutions, well-suited priors are required to reduce the solution space. In this paper, we analyze the characteristics of low-light images and their illumination component and identify a trivial solution not taken into consideration by the previous unsupervised state-of-the-art methods. The challenge comes from the fact that the trivial solution cannot be completely eliminated from the feasible set as it corresponds to the true solution when the low-light image contains a light source or an overexposed area. To address this issue, we propose a new regularization term which only remove absurd solutions and keep plausible ones in the set. To demonstrate the efficiency of the proposed prior, we conduct our experiments using deep image priors in a framework similar to the recent work RetinexDIP and an in-depth ablation study. Finally, we observe no more halo artefacts in the restored image. For all-but-one metrics, our unsupervised approach gives results as good as the supervised state-of-the-art indicating the potential of this framework for low-light image enhancement. Arthur Lecert, Renaud Fraisse, Aline Roumy, Christine Guillemot |
ICIP | 4 |
| 2022 | Statistical Analysis of Inter Coding in VVC Test Model (VTM)abstractThe promising compression efficiency improvement of Versatile Video Coding (VVC) compared to High Efficiency Video Coding (HEVC) [1] comes at the cost of a non-negligible encoder-side complexity. The largely increased complexity overhead is a possible obstacle towards its industrial implementation. Many papers have proposed acceleration methods for VVC. Still, a better understanding of VVC complexity, especially related to new partitions and coding tools, is desirable to help the design of new and better acceleration methods. For this purpose, statistical analyses have been conducted, with a focus on Coding Unit (CU) sizes and inter coding modes. Yiqun Liu 0013, Mohsen Abdoli, Thomas Guionnet, Christine Guillemot, Aline Roumy |
ICIP | 4 |
| 2022 | Omni-NeRF: Neural Radiance Field from 360° Image CapturesabstractThis paper tackles the problem of novel view synthesis (NVS) from 360° images with imperfect camera poses or intrinsic parameters. We propose a novel end-to-end framework for training Neural Radiance Field (NeRF) models given only 360° RGB images and their rough poses, which we refer to as Omni-NeRF. We extend the pinhole camera model of NeRF to a more general camera model that better fits omni-directional fish-eye lenses. The approach jointly learns the scene geometry and optimizes the camera parameters without knowing the fisheye projection. Thomas Maugey, Sebastian Knorr, Christine Guillemot |
ICME | 4 |
| 2022 | Quasi Lossless Satellite Image CompressionabstractWe describe an end-to-end trainable neural network for satel-lite image compression. The proposed approach builds upon an image compression scheme based on variational auto-encoders with a learned hyperprior that captures depen-dencies in the latent space for entropy coding. We explore this architecture in light of specificities of satellite imaging: processing constraints onboard the satellite (complexity and memory constraints) and quality needed in terms of reconstruction for the processing task on the ground. We explore data augmentation to improve the reconstruction of challenging image patterns. The proposed model outperforms the current standard of lossy image compression onboard satel-lite based on JPEG 2000, as well as the initial hyper-prior architecture designed for natural images. Pascal Bacchus, Renaud Fraisse, Aline Roumy, Christine Guillemot |
IGARSS | 4 |
| 2022 | Motion Compensation-based Low-Complexity Decoder Side Depth Estimation for MPEG Immersive VideoabstractDecoder-Side Depth Estimation (DSDE) is a system firstly enabled in the novel MPEG Immersive Video (MIV) coding standard. In DSDE, only texture components are coded, while the depth is estimated at the decoder-side. This is motivated by previous work, which has shown high coding gain and pixel rate savings in DSDE. However, the computational complexity remains a concern, as high quality depth search has a high runtime and memory requirement. In this work we extend the concept of depth estimation to depth recovery. Using this mode, the decoder-side depth information is recovered through motion compensation utilizing the displacement vectors contained in the texture bitstream. This strategy enables us to replace most of the complex depth estimation processes with a simple motion compensation step, a decision that is drawn on the encoder-side and signaled per coding unit. With only minor losses in terms of synthesis PSNR and similar perceptual quality in terms of MS-SSIM, the complexity is significantly reduced. Depending on the acceptable loss, up to 80 % of the moving objects depth may be motion compensated instead of estimated by a depth estimator translating into a speed-up of a factor of 104 for inter-frames compared to the reference depth estimator. Patrick Garus, Félix Henry, Thomas Maugey, Christine Guillemot |
MMSP | 4 |
| 2022 | Decoder Side Multiplane Images using Geometry Assistance SEI for MPEG Immersive VideoabstractThe MPEG Immersive Video (MIV) standard enables a novel technology denoted as decoder side depth estimation (DSDE) by introducing a dedicated Geometry Absent profile. In DSDE only texture information is coded and the corresponding geometry is reconstructed on the decoder side. MIV further enables the coding of side-information useful to the geometry reconstruction, denoted as Geometry Assistance SEI message. An emerging format for immersive video are Multiplane Images, which is investigated for feasibility in coding systems due to their promising rendering quality with complex sequences. In this work, we show that MIV can be used to construct block-based Multiplane Images on the decoder-side and to enhance the view synthesis performance utilizing the Geometry Assistance SEI. In a complexity-aware setting using only 32 planes, up to 6 dB of quality improvement is achieved compared to the reference. Patrick Garus, Félix Henry, Thomas Maugey, Christine Guillemot |
MMSP | 4 |
| 2022 | Axial refocusing precision model with light fields
Zhaolin Xiao, Jinglei Shi, Xiaoran Jiang, Christine Guillemot |
Signal Process. Image Commun. | 4 |
| 2022 | Immersive Video Coding: Should Geometry Information Be Transmitted as Depth Maps?abstractImmersive video often refers to multiple views with texture and scene geometry information, from which different viewports can be synthesized on the client side. To design efficient immersive video coding solutions, it is desirable to minimize bitrate, pixel rate and complexity. We investigate whether the classical approach of sending the geometry of a scene as depth maps is appropriate to serve this purpose. Previous work shows that bypassing depth transmission entirely and estimating depth at the client side improves the synthesis performance while saving bitrate and pixel rate. In order to understand if the encoder side depth maps contain information that is beneficial to be transmitted, we first explore a hybrid approach which enables partial depth map transmission using a block-based RD-based decision in the depth coding process. This approach reveals that partial depth map transmission may improve the rendering performance but does not present a good compromise in terms of compression efficiency. This led us to address the remaining drawbacks of decoder side depth estimation: complexity and depth map inaccuracy. We propose a novel system that takes advantage of high quality depth maps at the server side by encoding them into lightweight features that support the depth estimator at the client side. These features allow reducing the amount of data that has to be handled during decoder side depth estimation by 88%, which significantly speeds up the cost computation and the energy minimization of the depth estimator. Furthermore, −46.0% and −37.9% average synthesis BD-Rate gains are achieved compared to the classical approach with depth maps estimated at the encoder. Patrick Garus, Félix Henry, Joël Jung, Thomas Maugey, Christine Guillemot |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Advanced Scalability for Light Field Image Coding
Hadi Amirpour, Christine Guillemot, Mohammed Ghanbari 0001, Christian Timmerer |
IEEE Trans. Image Process. | 2 |
| 2022 | An Untrained Neural Network Prior for Light Field CompressionabstractDeep generative models have proven to be effective priors for solving a variety of image processing problems. However, the learning of realistic image priors, based on a large number of parameters, requires a large amount of training data. It has been shown recently, with the so-called deep image prior (DIP), that randomly initialized neural networks can act as good image priors without learning. In this paper, we propose a deep generative model for light fields, which is compact and which does not require any training data other than the light field itself. To show the potential of the proposed generative model, we develop a complete light field compression scheme with quantization-aware learning and entropy coding of the quantized weights. Experimental results show that the proposed method yields very competitive results compared with state-of-the-art light field compression methods, both in terms of PSNR and MS-SSIM metrics. Xiaoran Jiang, Jinglei Shi, Christine Guillemot |
IEEE Trans. Image Process. | 3 |
| 2022 | A Light Field FDL-HCGH Feature in Scale-Disparity SpaceabstractMany computer vision applications rely on feature detection and description, hence the need for computationally efficient and robust 4D light field (LF) feature detectors and descriptors. In this paper, we propose a novel light field feature descriptor based on the Fourier disparity layer representation, for light field imaging applications. After the Harris feature detection in a scale-disparity space, the proposed feature descriptor is then extracted using a circular neighborhood rather than a square neighborhood. It is shown to yield more accurate feature matching, compared with the LiFF LF feature, with a lower computational complexity. In order to evaluate the feature matching performance with the proposed descriptor, we generated a synthetic stereo LF dataset with ground truth matching points. Experimental results with synthetic and real-world dataset show that our solution outperforms existing methods in terms of both feature detection robustness and feature matching accuracy. Haiyan Jin, Zhaolin Xiao, Christine Guillemot |
IEEE Trans. Image Process. | 4 |
| 2022 | A Sparse Non-parametric BRDF ModelabstractThis paper presents a novel sparse non-parametric Bidirectional Reflectance Distribution Function (BRDF) model derived using a machine learning approach to represent the space of possible BRDFs using a set of multidimensional sub-spaces, or dictionaries. By training the dictionaries under a sparsity constraint, the model guarantees high-quality representations with minimal storage requirements and an inherent clustering of the BDRF-space. The model can be trained once and then reused to represent a wide variety of measured BRDFs. Moreover, the proposed method is flexible to incorporate new unobserved data sets, parameterizations, and transformations. In addition, we show that any two, or more, BRDFs can be smoothly interpolated in the coefficient space of the model rather than the significantly higher-dimensional BRDF space. The proposed sparse BRDF model is evaluated using the MERL, DTU, and RGL-EPFL BRDF databases. Experimental results show that the proposed approach results in about 9.75dB higher signal-to-noise ratio on average for rendered images as compared to current state-of-the-art models. Tanaboon Tongbuasirilai, Jonas Unger, Christine Guillemot, Ehsan Miandji |
ACM Trans. Graph. | 3 |
| 2021 | A Light Field FDL-HSIFT Feature in Scale-Disparity SpaceabstractMany computer vision applications rely on feature matching, hence the need for computationally efficient and robust 4D light field (LF) feature detectors and descriptors for applications using this imaging modality. In this paper, we propose a novel LF feature extraction method in the scale-disparity space, based on a Fourier disparity layer representation. The proposed feature extraction takes advantage of both the Harris feature detector and SIFT descriptor, and is shown to yield more accurate feature matching, compared with the LiFF light field feature with low computational complexity. In order to evaluate the feature matching performance with the proposed descriptor, we generated synthetic LF datasets with ground truth matching points. Experimental results with synthetic and real datasets show that, our solution outperforms existing methods in terms of both feature detection robustness and feature matching accuracy. Zhaolin Xiao, Haiyan Jin, Christine Guillemot |
ICIP | 4 |
| 2021 | Analysis of Top-Down Connections in Multi-Layered Convolutional Sparse CodingabstractConvolutional Neural Networks (CNNs) have been instrumental in the recent advances in machine learning, with applications to media applications. Multi-Layered Convolutional Sparse Coding (ML-CSC) based on a cascade of convolutional layers in which each layer can be approximately explained by the following layer can be seen as a biologically inspired framework. However, both CNNs and ML-CSC networks lack top-down information flows that are studied in neuroscience for understanding the mechanisms of the mammal cortex. A successful implementation of such top-down connections could lead to another leap in machine learning and media applications. This study analyses the effects of a feedback connection on an ML-CSC network, considering trade-off between sparsity and reconstruction error, support recovery rate, and mutual coherence in trained dictionaries. We find that using the feedback connection during training impacts the mutual coherence of the dictionary in a way that the equivalence between the l0-and l1-norm is verified for a smaller range of sparsity values. Experimental results show that the use of feedback during training does not favour inference with feedback, in terms of sparse support recovery rates. However, when the sparsity constraints are given a lower weight, the use of feedback at inference time is beneficial, in terms of support recovery rates. Joakim Edlund, Christine Guillemot, Mårten Sjöström |
MMSP | 2 |
| 2021 | Regularizing the Deep Image Prior with a Learned Denoiser for Linear Inverse ProblemsabstractWe propose an optimization method coupling a learned denoiser with the untrained generative model, called deep image prior (DIP) in the framework of the Alternating Direction Method of Multipliers (ADMM) method. We also study different regularizers of DIP optimization, for inverse problems in imaging, focusing in particular on denoising and super-resolution. The goal is to make the best of the untrained DIP and of a generic regularizer learned in a supervised manner from a large collection of images. When placed in the ADMM framework, the denoiser is used as a proximal operator and can be learned independently of the considered inverse problem. We show the benefits of the proposed method, in comparison with other regularized DIP methods, for two linear inverse problems, i.e., denoising and super-resolution. Rita Fermanian, Mikael Le Pendu, Christine Guillemot |
MMSP | 3 |
| 2021 | A learning-based view extrapolation method for axial super-resolution
Zhaolin Xiao, Jinglei Shi, Xiaoran Jiang, Christine Guillemot |
Neurocomputing | 4 |
| 2021 | A Lightweight Neural Network for Monocular View Generation With Occlusion HandlingabstractIn this article, we present a very lightweight neural network architecture, trained on stereo data pairs, which performs view synthesis from one single image. With the growing success of multi-view formats, this problem is indeed increasingly relevant. The network returns a prediction built from disparity estimation, which fills in wrongly predicted regions using a occlusion handling technique. To do so, during training, the network learns to estimate the left-right consistency structural constraint on the pair of stereo input images, to be able to replicate it at test time from one single image. The method is built upon the idea of blending two predictions: a prediction based on disparity estimation and a prediction based on direct minimization in occluded regions. The network is also able to identify these occluded areas at training and at test time by checking the pixelwise left-right consistency of the produced disparity maps. At test time, the approach can thus generate a left-side and a right-side view from one input image, as well as a depth map and a pixelwise confidence measure in the prediction. The work outperforms visually and metric-wise state-of-the-art approaches on the challenging KITTI dataset, all while reducing by a very significant order of magnitude (5 or 10 times) the required number of parameters (6.5 M). Simon Evain, Christine Guillemot |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Compressively sampled light field reconstruction using orthogonal frequency selection and refinement
Fatma Hawary, Guillaume Boisson, Christine Guillemot, Philippe Guillotel |
Signal Process. Image Commun. | 3 |
| 2021 | Rate-Distortion Optimized Graph Coarsening and Partitioning for Light Field CodingabstractGraph-based transforms are powerful tools for signal representation and energy compaction. However, their use for high dimensional signals such as light fields poses obvious problems of complexity. To overcome this difficulty, one can consider local graph transforms defined on supports of limited dimension, which may however not allow us to fully exploit long-term signal correlation. In this paper, we present methods to optimize local graph supports in a rate distortion sense for efficient light field compression. A large graph support can be well adapted for compression efficiency, however at the expense of high complexity. In this case, we use graph reduction techniques to make the graph transform feasible. We also consider spectral clustering to reduce the dimension of the graph supports while controlling both rate and complexity. We derive the distortion and rate models which are then used to guide the graph optimization. We describe a complete light field coding scheme based on the proposed graph optimization tools. Experimental results show rate-distortion performance gains compared to the use of fixed graph support. The method also provides competitive results when compared against HEVC-based and the JPEG Pleno light field coding schemes. We also assess the method against a homography-based low rank approximation and a Fourier disparity layer based coding method. Mira Rizkallah, Thomas Maugey, Christine Guillemot |
IEEE Trans. Image Process. | 3 |
| 2020 | Learning Fused Pixel and Feature-Based View Reconstructions for Light FieldsabstractIn this paper, we present a learning-based framework for light field view synthesis from a subset of input views. Building upon a light-weight optical flow estimation network to obtain depth maps, our method employs two reconstruction modules in pixel and feature domains respectively. For the pixel-wise reconstruction, occlusions are explicitly handled by a disparity-dependent interpolation filter, whereas inpainting on disoccluded areas is learned by convolutional layers. Due to disparity inconsistencies, the pixel-based reconstruction may lead to blurriness in highly textured areas as well as on object contours. On the contrary, the feature-based reconstruction well performs on high frequencies, making the reconstruction in the two domains complementary. End-to-end learning is finally performed including a fusion module merging pixel and feature-based reconstructions. Experimental results show that our method achieves state-of-the-art performance on both synthetic and real-world datasets, moreover, it is even able to extend light fields' baseline by extrapolating high quality views without additional training. Jinglei Shi, Xiaoran Jiang, Christine Guillemot |
CVPR | 3 |
| 2020 | Colour Compression of Plenoptic Point Clouds Using Raht-Klt with Prior Colour Clustering and Specular/Diffuse Component SeparationabstractThe recently introduced plenoptic point cloud representation marries a 3D point cloud with a light field. Instead of each point being associated with a single colour value, there can be multiple values to represent the colour at that point as perceived from different viewpoints. This representation was introduced together with a compression technique for the multi-view colour vectors, which is an extension of the RAHT method for point cloud attribute coding. In the current paper, we demonstrate that the best-proposed RAHT extension, RAHT-KLT, can be improved by performing a prior subdivision of the plenoptic point cloud into clusters based on similar colour values, followed by a separation of each cluster into specular and diffuse components, and coding each component separately with RAHT-KLT. Our proposed improvements are shown to achieve better rate-distortion results than the original RAHT-KLT method. Maja Krivokuca, Christine Guillemot |
ICASSP | 2 |
| 2020 | Color and Angular Reconstruction of Light Fields from Incomplete-Color Coded ProjectionsabstractWe present a simple variational approach for reconstructing color light fields (LFs) in the compressed sensing (CS) framework with very low sampling ratio, using both coded masks and color filter arrays (CFAs). A coded mask is placed in front of the camera sensor to optically modulate incoming rays, while a CFA is assumed to be implemented at the sensor level to compress color information. Hence, the LF coded projections, operated by a combination of the coded mask and the CFA, measure incomplete color samples with a three-times-lower sampling ratio than reference methods that assume full color (channel-by-channel) acquisition. We then derive adaptive algorithms to directly reconstruct the light field from raw sensor measurements by minimizing a convex energy composed of two terms. The first one is the data fidelity term which takes into account the use of CFAs in the imaging model, and the second one is a regularization term which favors the sparse representation of light fields in a specific transform domain. Experimental results show that the proposed approach produces a better reconstruction both in terms of visual quality and quantitative performance when compared to reference reconstruction methods that implicitly assume prior color interpolation of coded projections. Christine Guillemot |
ICASSP | 2 |
| 2020 | Sub-Dip: Optimization On A Subspace With Deep Image Prior Regularization And Application To SuperresolutionabstractThe Deep Image Prior has been recently introduced to solve inverse problems in image processing with no need for training data other than the image itself. However, the original training algorithm of the Deep Image Prior constrains the reconstructed image to be on a manifold described by a convolutional neural network. For some problems, this neglects prior knowledge and can render certain regularizers ineffective. This work proposes an alternative approach that relaxes this constraint and fully exploits all prior knowledge. We evaluate our algorithm on the problem of reconstructing a high-resolution image from a downsampled version and observe a significant improvement over the original Deep Image Prior algorithm. Alexander Sagel, Aline Roumy, Christine Guillemot |
ICASSP | 3 |
| 2020 | Angularly Consistent Light Field Video InterpolationabstractIn this paper, we address the problem of temporal interpolation of sparsely sampled video light fields using dense scene flows. Given light fields at two time instants, the goal is to interpolate an intermediate light field to form a spatially, angularly and temporally coherent light field video sequence. We first compute angularly coherent bidirectional scene flows between the two input light fields. We then use the optical flows and the two light fields as inputs to a convolutional neural network that synthesizes independently the views of the light field at an intermediate time. In order to measure the angular consistency of a light field, we propose a new metric based on epipolar geometry. Experimental results show that the proposed method produces light fields that are angularly coherent while keeping similar temporal and spatial consistency as state-of-the-art video frame interpolation methods. Pierre David 0001, Mikael Le Pendu, Christine Guillemot |
ICME | 3 |
| 2020 | Sphere Mapping for Feature Extraction From 360° Fish-Eye CapturesabstractEquirectangular projection is commonly used to map 360° captures into planar representation, so that existent processing methods can be directly applied to such content. Such format introduces stitching distortions that could impact the efficiency of further processing such as camera pose estimation, 3D point localization and depth estimation. Indeed, even if some algorithms, mainly feature descriptors, tend to remap the projected images into a sphere, important radial distortions remain existent in the processed data. In this paper, we propose to adapt the spherical model to the geometry of the 360° fish-eye camera, and avoid the stitching process. We consider the angular coordinates of feature points on the sphere for evaluation. We assess the precision of different operations such as camera rotation angle estimation and 3D point depth calculation on spherical camera images. Experimental results show that the proposed fish-eye adapted sphere mapping allows more stability in angle estimation, as well as in 3D point localization, compared to the one on projected and stitched contents. Fatma Hawary, Thomas Maugey, Christine Guillemot |
MMSP | 3 |
| 2020 | Single Sensor Compressive Light Field Video CameraabstractAbstract This paper presents a novel compressed sensing (CS) algorithm and camera design for light field video capture using a single sensor consumer camera module. Unlike microlens light field cameras which sacrifice spatial resolution to obtain angular information, our CS approach is designed for capturing light field videos with high angular, spatial, and temporal resolution. The compressive measurements required by CS are obtained using a random color‐coded mask placed between the sensor and aperture planes. The convolution of the incoming light rays from different angles with the mask results in a single image on the sensor; hence, achieving a significant reduction on the required bandwidth for capturing light field videos. We propose to change the random pattern on the spectral mask between each consecutive frame in a video sequence and extracting spatio‐angular‐spectral‐temporal 6D patches. Our CS reconstruction algorithm for light field videos recovers each frame while taking into account the neighboring frames to achieve significantly higher reconstruction quality with reduced temporal incoherencies, as compared with previous methods. Moreover, a thorough analysis of various sensing models for compressive light field video acquisition is conducted to highlight the advantages of our method. The results show a clear advantage of our method for monochrome sensors, as well as sensors with color filter arrays. Saghi Hajisharif, Ehsan Miandji, Christine Guillemot, Jonas Unger |
Comput. Graph. Forum | 3 |
| 2020 | Light Field Super-Resolution Using a Low-Rank Prior and Deep Convolutional Neural NetworksabstractLight field imaging has recently known a regain of interest due to the availability of practical light field capturing systems that offer a wide range of applications in the field of computer vision. However, capturing high-resolution light fields remains technologically challenging since the increase in angular resolution is often accompanied by a significant reduction in spatial resolution. This paper describes a learning-based spatial light field super-resolution method that allows the restoration of the entire light field with consistency across all angular views. The algorithm first uses optical flow to align the light field and then reduces its angular dimension using low-rank approximation. We then consider the linearly independent columns of the resulting low-rank model as an embedding, which is restored using a deep convolutional neural network (DCNN). The super-resolved embedding is then used to reconstruct the remaining views. The original disparities are restored using inverse warping where missing pixels are approximated using a novel light field inpainting algorithm. Experimental results show that the proposed method outperforms existing light field super-resolution algorithms, achieving PSNR gains of 0.23 dB over the second best performing method. The performance is shown to be further improved using iterative back-projection as a post-processing step. Reuben A. Farrugia, Christine Guillemot |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | A simple framework to leverage state-of-the-art single-image super-resolution methods to restore light fields
Reuben A. Farrugia, Christine Guillemot |
Signal Process. Image Commun. | 2 |
| 2020 | Local Low Rank Approximation With a Parametric Disparity Model for Light Field CompressionabstractWe address the problem of light field dimensionality reduction for compression. We describe a local low rank approximation method using a parametric disparity model. The local support of the approximation is defined by super-rays. A super-ray can be seen as a set of super-pixels that are coherent across all light field views. A dedicated super-ray construction method is first described that constrains the super-pixels forming a given super-ray to be all of the same shape and size, dealing with occlusions. This constraint is needed so that the super-rays can be used as supports of angular dimensionality reduction based on low rank matrix approximation. The light field low rank assumption depends on how much the views are correlated, i.e. on how well they can be aligned by disparity compensation. We first introduce a parametric model describing the local variations of disparity within each super-ray. We then consider two methods for estimating the model parameters. The first method simply fits the model on an input disparity map. We then introduce a disparity estimation method using a low rank prior. This method alternatively searches for the best parameters of the disparity model and of the low rank approximation. We assess the proposed disparity parametric model, first assuming that the disparity is constant within a super-ray, and second by considering an affine disparity model. We show that using the proposed disparity parametric model and estimation algorithm gives an alignment of super-pixels across views that favours the low rank approximation compared with using disparity estimated with classical computer vision methods. The low rank matrix approximation is computed on the disparity compensated super-rays using a singular value decomposition (SVD). A coding algorithm is then described for the different components of the proposed disparity-compensated low rank approximation. Experimental results show performance gains, with a rate saving going up to 92.61%, compared with the JPEG Pleno anchor, for real light fields captured by a Lytro Illum camera. The rate saving goes up to 37.72% with synthetic light fields. The approach is also shown to outperform an HEVC-based light field compression scheme. Elian Dib, Mikael Le Pendu, Xiaoran Jiang, Christine Guillemot |
IEEE Trans. Image Process. | 4 |
| 2020 | Context-Adaptive Neural Network-Based Prediction for Image CompressionabstractThis paper describes a set of neural network architectures, called Prediction Neural Networks Set (PNNS), based on both fully-connected and convolutional neural networks, for intra image prediction. The choice of neural network for predicting a given image block depends on the block size, hence does not need to be signalled to the decoder. It is shown that, while fully-connected neural networks give good performance for small block sizes, convolutional neural networks provide better predictions in large blocks with complex textures. Thanks to the use of masks of random sizes during training, the neural networks of PNNS well adapt to the available context that may vary, depending on the position of the image block to be predicted. When integrating PNNS into a H.265 codec, PSNRrate performance gains going from 1:46% to 5:20% are obtained. These gains are on average 0:99% larger than those of prior neural network based methods. Unlike the H.265 intra prediction modes, which are each specialized in predicting a specific texture, the proposed PNNS can model a large set of complex textures. Thierry Dumas, Aline Roumy, Christine Guillemot |
IEEE Trans. Image Process. | 3 |
| 2020 | Optical-Flow Based Nonlinear Weighted Prediction for SDR and Backward Compatible HDR Video CodingabstractTone Mapping Operators (TMO) designed for videos can be classified into two categories. In a first approach, TMOs are temporal filtered to reduce temporal artifacts and provide a Standard Dynamic Range (SDR) content with improved temporal consistency. This however does not improve the SDR coding Rate Distortion (RD) performances. A second approach is to design the TMO with the goal of optimizing the SDR coding rate-distortion performances. This second category of methods may lead to SDR videos altering the artistic intent compared with the produced HDR content. In this paper, we combine the benefits of the two approaches by introducing new Weighted Prediction (WP) methods inside the HEVC SDR codec. As a first step, we demonstrate the interest of the WP methods compared to TMO optimized for RD performances. Then we present the newly introduced WP algorithm and WP modes. The WP algorithm consists in performing a global motion compensation between frames using an optical flow, and the new modes are based on non linear functions in contrast with the literature using only linear functions. The contribution of each novelty is studied independently and in a second time they are all put in competition to maximize the RD performances. Tests were made for HDR backward compatible compression but also for SDR compression only. In both cases, the proposed WP methods improve the RD performances while maintaining the SDR temporal coherency. David Gommelet, Julien Le Tanou, Aline Roumy, Michaël Ropert, Christine Guillemot |
IEEE Trans. Image Process. | 5 |
| 2020 | Geometry-Aware Graph Transforms for Light Field Compact RepresentationabstractThe paper addresses the problem of energy compaction of dense 4D light fields by designing geometry-aware local graph-based transforms. Local graphs are constructed on super-rays that can be seen as a grouping of spatially and geometry-dependent angularly correlated pixels. Both non separable and separable transforms are considered. Despite the local support of limited size defined by the super-rays, the Laplacian matrix of the non separable graph remains of high dimension and its diagonalization to compute the transform eigen vectors remains computationally expensive. To solve this problem, we then perform the local spatio-angular transform in a separable manner. We show that when the shape of corresponding super-pixels in the different views is not isometric, the basis functions of the spatial transforms are not coherent, resulting in decreased correlation between spatial transform coefficients. We hence propose a novel transform optimization method that aims at preserving angular correlation even when the shapes of the super-pixels are not isometric. Experimental results show the benefit of the approach in terms of energy compaction. A coding scheme is also described to assess the rate-distortion perfomances of the proposed transforms and is compared to state of the art encoders namely HEVC-lozenge [1], JPEG pleno 1.1 [2], HEVC-pseudo [3] and HLRA [4]. Mira Rizkallah, Xin Su 0003, Thomas Maugey, Christine Guillemot |
IEEE Trans. Image Process. | 4 |
| 2020 | Prediction and Sampling With Local Graph Transforms for Quasi-Lossless Light Field CompressionabstractGraph-based transforms have been shown to be powerful tools in terms of image energy compaction. However, when the size of the support increases to best capture signal dependencies, the computation of the basis functions becomes rapidly untractable. This problem is in particular compelling for high dimensional imaging data such as light fields. The use of local transforms with limited supports is a way to cope with this computational difficulty. Unfortunately, the locality of the support may not allow us to fully exploit long term signal dependencies present in both the spatial and angular dimensions of light fields. This paper describes sampling and prediction schemes with local graph-based transforms enabling to efficiently compact the signal energy and exploit dependencies beyond the local graph support. The proposed approach is investigated and is shown to be very efficient in the context of spatio-angular transforms for quasi-lossless compression of light fields. Mira Rizkallah, Thomas Maugey, Christine Guillemot |
IEEE Trans. Image Process. | 3 |
| 2019 | Deep Frame Interpolation for Video CompressionabstractDeep neural networks have been recently proposed to solve video interpolation tasks. Given a past and future frame, such networks can be trained to successfully predict the intermediate frame(s). In the context of video compression, these architectures could be useful as an additional inter-prediction mode. Current inter-prediction methods rely on block-matching techniques to estimate the motion between consecutive frames. This approach has severe limitations for handling complex non-translational motions, and is still limited to block-based motion vectors. This paper presents a deep frame interpolation network for video compression aiming at solving the previous limitations, i.e. able to cope with all types of geometrical deformations by providing a dense motion compensation. Experiments with the classical bi-directional hierarchical video coding structure demonstrate the efficiency of the proposed approach over the traditional tools of the HEVC codec. Jean Bégaint, Franck Galpin, Philippe Guillotel, Christine Guillemot |
DCC | 4 |
| 2019 | Super-Ray Based Low Rank Approximation for Light Field CompressionabstractWe describe a local low rank approximation method based on super-rays for light field compression. Super-rays can be seen as a set of super-pixels that are coherent across all light field views. A super-ray based disparity estimation method is proposed using a low rank prior, in order to be able to align all the super-pixels forming each super-ray. A dedicated super-ray construction method is described that constrains the super-pixels forming a given super-ray to be all of the same shape and size, dealing with occlusions. This constraint is needed so that the super-rays can be used as a support of angular dimensionality reduction based on low rank matrix approximation. A low rank matrix approximation is then computed on the disparity compensated super-rays using a singular value decomposition (SVD). A coding algorithm is then described for the different components of the resulting low rank approximation. Experimental results show performance gains compared with two reference light field coding schemes (HEVC-based scheme and JPEG-Pleno VM 1.1). Elian Dib, Mikael Le Pendu, Xiaoran Jiang, Christine Guillemot |
DCC | 4 |
| 2019 | Graph-Based Spatio-Angular Prediction for Quasi-Lossless Compression of Light FieldsabstractGraph-based transforms have been shown to be powerful tools for image compression. However, the computation of the basis functions becomes rapidly untractable when the support increases, i.e. when the dimension of the data is high as in the case of light fields. Local transforms with limited supports have been investigated to cope with this difficulty. Nevertheless, the locality of the support may not allow us to fully exploit long term dependencies in the signal. In this paper, we describe a graph based prediction solution that allows taking advantage of intra prediction mechanisms as well as of the good energy compaction properties of the graph transform. The approach relies on a separable spatio-angular transform and derives low frequency spatio-angular coefficients from one single compressed reference view and from the high angular frequency coefficients. In the tests, we used HEVC-Intra, with QP=0, to encode the reference frame with high quality. The high angular frequency coefficients containing very little energy are coded using a simple entropy coder. The approach is shown to be very efficient in a context of high quality quasi-lossless compression of light fields. Mira Rizkallah, Thomas Maugey, Christine Guillemot |
DCC | 3 |
| 2019 | Light Field Denoising Using 4D Anisotropic DiffusionabstractIn this paper, we present a novel light field denoising algorithm using a vector-valued regularization operating in the 4D ray space. More precisely, the method performs a PDE-based anisotropic diffusion along directions defined by local structures in the 4D ray space. It does not require prior estimation of disparity maps. The local structures in the 4D light field are extracted using a 4D tensor structure. The paper then describes the strategy retained for setting the diffusion tensor parameters for the targeted denoising application. It then analyzes the influence of the model parameters on the denoising performance. Experimental results show that the proposed denoising algorithm performs well compared to state of the art methods while keeping tractable complexity. Pierre Allain, Laurent Guillo, Christine Guillemot |
ICASSP | 3 |
| 2019 | Denoising of 3D Point Clouds Constructed from Light FieldsabstractLight fields are 4D signals capturing rich information from a scene. The availability of multiple views enables scene depth estimation, that can be used to generate 3D point clouds. The constructed 3D point clouds, however, generally contain distortions and artefacts primarily caused by inaccuracies in the depth maps. This paper describes a method for noise removal in 3D point clouds constructed from light fields. While existing methods discard outliers, the proposed approach instead attempts to correct the positions of points, and thus reduce noise without removing any points, by exploiting the consistency among views in a light-field. The proposed 3D point cloud construction and denoising method exploits uncertainty measures on depth values. We also investigate the possible use of the corrected point cloud to improve the quality of the depth maps estimated from the light field. Chris Galea, Christine Guillemot |
ICASSP | 2 |
| 2019 | A Learning Based Depth Estimation Framework for 4D Densely and Sparsely Sampled Light FieldsabstractThis paper proposes a learning based solution to disparity (depth) estimation for either densely or sparsely sampled light fields. Disparity between stereo pairs among a sparse subset of anchor views is first estimated by a fine-tuned FlowNet 2.0 network adapted to disparity prediction task. These coarse estimates are fused by exploiting the photo-consistency warping error, and refined by a Multi-view Stereo Refinement Network (MSRNet). The propagation of disparity from anchor viewpoints towards other viewpoints is performed by an occlusion-aware soft 3D reconstruction method. The experiments show that, both for dense and sparse light fields, our algorithm outperforms significantly the state-of-the-art algorithms, especially for subpixel accuracy. Xiaoran Jiang, Jinglei Shi, Christine Guillemot |
ICASSP | 3 |
| 2019 | Sparse to Dense Scene Flow Estimation From Light FieldsabstractThe paper addresses the problem of scene flow estimation from sparsely sampled video light fields. The scene flow estimation method is based on an affine model in the 4D ray space that allows us to estimate a dense flow from sparse estimates in 4D clusters. A dataset of synthetic video light fields created for assessing scene flow estimation techniques is also described. Experiments show that the proposed method gives error rates on the optical flow components that are comparable to those obtained with state of the art optical flow estimation methods, while computing a more accurate disparity variation when compared with prior scene flow estimation techniques. Pierre David 0001, Mikael Le Pendu, Christine Guillemot |
ICIP | 3 |
| 2019 | Light Field Compression Using Fourier Disparity LayersabstractIn this paper, we present a compression method for light fields based on the Fourier Disparity Layer representation. This light field representation consists in a set of layers that can be efficiently constructed in the Fourier domain from a sparse set of views, and then used to reconstruct intermediate viewpoints without requiring a disparity map. In the proposed compression scheme, a subset of light field views is encoded first and used to construct a Fourier Disparity Layer model from which a second subset of views is predicted. After encoding and decoding the residual of those predicted views, a larger set of decoded views is available, allowing us to refine the layer model in order to predict the next views with increased accuracy. The procedure is repeated until the complete set of light field views is encoded. Following this principle, we investigate in the paper different scanning orders of the light field views and analyse their respective efficiencies regarding the compression performance. Elian Dib, Mikael Le Pendu, Christine Guillemot |
ICIP | 3 |
| 2019 | Long Short Term Memory Networks for Light Field View SynthesisabstractBecause light field devices have a limited angular resolution, artificially reconstructing intermediate views is an interesting task. In this work, we propose a novel way to solve this problem using deep learning. In particular, the use of Long Short Term Memory Networks on a plane sweep volume is proposed. The approach has the advantage of having very few parameters and can be run on sequences with arbitrary length. We show that our approach yields results that are competitive with the state-of-the-art for dense light fields. Experimental results also show promising results with light fields with wider baselines. Matthieu Hog, Neus Sabater, Christine Guillemot |
ICIP | 3 |
| 2019 | Bypassing Depth Maps Transmission For Immersive Video CodingabstractThis paper addresses several downsides of the system under development in MPEG-I for coding and transmission of immersive media. We present a solution, which enables Depth-Image-Based Rendering for immersive video applications, while lifting the requirement of transmitting depth information. Instead, we estimate the depth information on the client-side from the transmitted views. The approach leads to an impressive rate saving (37.3% in average). Preserving perceptual quality in terms of MS-SSIM of synthesized views, it yields to 24.6% rate reduction for the same quality of reconstructed views after residue transmission under the MPEG-I common test conditions. Simultaneously, the required pixel rate, i.e. the number of pixels processed per second by the decoder, is reduced by 50% for any test sequence. To the author's knowledge, this is the first time that such an approach is under consideration in the context of immersive video coding. Patrick Garus, Joël Jung, Thomas Maugey, Christine Guillemot |
PCS | 4 |
| 2019 | A Fourier Disparity Layer Representation for Light FieldsabstractIn this paper, we present a new Light Field representation for efficient Light Field processing and rendering called Fourier Disparity Layers (FDL). The proposed FDL representation samples the Light Field in the depth (or equivalently the disparity) dimension by decomposing the scene as a discrete sum of layers. The layers can be constructed from various types of Light Field inputs, including a set of sub-aperture images, a focal stack, or even a combination of both. From our derivations in the Fourier domain, the layers are simply obtained by a regularized least square regression performed independently at each spatial frequency, which is efficiently parallelized in a GPU implementation. Our model is also used to derive a gradient descent-based calibration step that estimates the input view positions and an optimal set of disparity values required for the layer construction. Once the layers are known, they can be simply shifted and filtered to produce different viewpoints of the scene while controlling the focus and simulating a camera aperture of arbitrary shape and size. Our implementation in the Fourier domain allows real-time Light Field rendering. Finally, direct applications such as view interpolation or extrapolation and denoising are presented and evaluated. Mikael Le Pendu, Christine Guillemot, Aljoscha Smolic |
IEEE Trans. Image Process. | 2 |
| 2019 | A Framework for Learning Depth From a Flexible Subset of Dense and Sparse Light Field ViewsabstractIn this paper, we propose a learning-based depth estimation framework suitable for both densely and sparsely sampled light fields. The proposed framework consists of three processing steps: initial depth estimation, fusion with occlusion handling, and refinement. The estimation can be performed from a flexible subset of input views. The fusion of initial disparity estimates, relying on two warping error measures, allows us to have an accurate estimation in occluded regions and along the contours. In contrast with methods relying on the computation of cost volumes, the proposed approach does not need any prior information on the disparity range. Experimental results show that the proposed method outperforms state-of-the-art light fields depth estimation methods, including prior methods based on deep neural architectures. Jinglei Shi, Xiaoran Jiang, Christine Guillemot |
IEEE Trans. Image Process. | 3 |
| 2018 | Fast Light Field Inpainting Propagation Using Angular Warping and Color-Guided Disparity Interpolation
Pierre Allain, Laurent Guillo, Christine Guillemot |
ACIVS | 3 |
| 2018 | Dynamic Super-Rays for Efficient Light Field Video Processing
Matthieu Hog, Neus Sabater, Christine Guillemot |
BMVC | 3 |
| 2018 | Autoencoder Based Image Compression: Can the Learning be Quantization Independent?abstractThis paper explores the problem of learning transforms for image compression via autoencoders. Usually, the rate-distortion performances of image compression are tuned by varying the quantization step size. In the case of autoencoders, this in principle would require learning one transform per rate-distortion point at a given quantization step size. Here, we show that comparable performances can be obtained with a unique learned transform. The different rate-distortion points are then reached by varying the quantization step size at test time. This approach saves a lot of training time. Thierry Dumas, Aline Roumy, Christine Guillemot |
ICASSP | 3 |
| 2018 | Graph-based Transforms for Predictive Light Field Compression based on Super-PixelsabstractIn this paper, we explore the use of graph-based transforms to capture correlation in light fields. We consider a scheme in which view synthesis is used as a first step to exploit inter-view correlation. Local graph-based transforms (GT) are then considered for energy compaction of the residue signals. The structure of the local graphs is derived from a coherent super-pixel over-segmentation of the different views. The GT is computed and applied in a separable manner with a first spatial unweighted transform followed by an inter-view GT. For the inter-view GT, both unweighted and weighted GT have been considered. The use of separable instead of non separable transforms allows us to limit the complexity inherent to the computation of the basis functions. A dedicated simple coding scheme is then described for the proposed GT based light field decomposition. Experimental results show a significant improvement with our method compared to the CNN view synthesis method and to the HEVC direct coding of the light field views. Mira Rizkallah, Xin Su 0003, Thomas Maugey, Christine Guillemot |
ICASSP | 4 |
| 2018 | Compressive 4D Light Field Reconstruction Using Orthogonal Frequency SelectionabstractWe present a new method for reconstructing a 4D light field from a random set of measurements. A 4D light field block can be represented by a sparse model in the Fourier domain. As such, the proposed algorithm reconstructs the light field, block by block, by selecting frequencies of the model that best fits the available samples, while enforcing orthogonality with the approximation residue. The method achieves a very high reconstruction quality, in terms of Peak Signal-to-Noise Ratio (PSNR). Experiments on several datasets show significant quality improvements of more than 1dB compared to state-of-the-art algorithms. Fatma Hawary, Guillaume Boisson, Christine Guillemot, Philippe Guillotel |
ICIP | 3 |
| 2018 | Depth Estimation with Occlusion Handling from a Sparse Set of Light Field ViewsabstractInternational audience Xiaoran Jiang, Mikael Le Pendu, Christine Guillemot |
ICIP | 3 |
| 2018 | High Dynamic Range Light Fields via Weighted Low Rank ApproximationabstractIn this paper, we propose a method for capturing High Dynamic Range (HDR) light fields with dense viewpoint sampling. Analogously to the traditional HDR acquisition process, several light fields are captured at varying exposures with a plenoptic camera. The RAW data is de-multiplexed to retrieve all light field viewpoints for each exposure and perform a soft detection of saturated pixels. Considering a matrix which concatenates all the vectorized views, we formulate the problem of recovering saturated areas as a Weighted Low Rank Approximation (WLRA) where the weights are defined from the soft saturation detection. We show that our algorithm successfully recovers the parallax in the over-exposed areas while the Truncated Nuclear Norm (TNN) minimization, traditionally used for single view HDR imaging, does not generalize to light fields. Advantages of our weighted approach as well as the simultaneous processing of all the viewpoints are also demonstrated in our experiments. Mikael Le Pendu, Christine Guillemot, Aljoscha Smolic |
ICIP | 2 |
| 2018 | Region-based models for motion compensation in video compressionabstractVideo codecs are primarily designed assuming that rigid, block-based, two-dimensional displacements are suitable models to describe the motion taking place in a scene. However, translational models are not sufficient to handle real world motion types such as camera zoom, shake, pan, shearing or changes in aspect ratio. We present here a region-based interprediction scheme to compensate such motion. The proposed mode is able to estimate multiple homography models in order to predict complex scene motion. We also introduce an affine photometric correction to each geometric model. Experiments on targeted sequences with complex motion demonstrate the efficiency of the proposed approach compared to the state-of-the-art HEVC video codec. Jean Bégaint, Franck Galpin, Philippe Guillotel, Christine Guillemot |
PCS | 4 |
| 2018 | A new framework for optimal facial landmark localization on light-field imagesabstractThe paper explores how light fields captured by plenoptic cameras can increase the performance of face landmark detection. The idea is to exploit light fields geometrical constraints to correct the position of points detected by classical face landmark detectors. These geometric constraints are used to enforce landmark points angular coherency across the different views of the light field, and by doing so to correct the positions of the landmarks on all views. The corrected landmark points are compared with ground-truth manual annotations of a set of 400 images corresponding to the central views of 400 light fields of faces with different pose and expression. Chiara Galdi, Lara Younes, Christine Guillemot, Jean-Luc Dugelay |
VCIP | 3 |
| 2018 | Region-Based Prediction for Image Compression in the CloudabstractThanks to the increasing number of images stored in the cloud, external image similarities can be leveraged to efficiently compress images by exploiting inter-images correlations. In this paper, we propose a novel image prediction scheme for cloud storage. Unlike current state-of-the-art methods, we use a semi-local approach to exploit inter-image correlation. The reference image is first segmented into multiple planar regions determined from matched local features and super-pixels. The geometric and photometric disparities between the matched regions of the reference image and the current image are then compensated. Finally, multiple references are generated from the estimated compensation models and organized in a pseudo-sequence to differentially encode the input image using classical video coding tools. Experimental results demonstrate that the proposed approach yields significant rate-distortion performance improvements compared with the current image inter-coding solutions such as high efficiency video coding. Jean Bégaint, Dominique Thoreau, Philippe Guillotel, Christine Guillemot |
IEEE Trans. Image Process. | 4 |
| 2018 | Light Field Inpainting Propagation via Low Rank Matrix CompletionabstractBuilding up on the advances in low rank matrix completion, this article presents a novel method for propagating the inpainting of the central view of a light field to all the other views. After generating a set of warped versions of the inpainted central view with random homographies, both the original light field views and the warped ones are vectorized and concatenated into a matrix. Because of the redundancy between the views, the matrix satisfies a low rank assumption enabling us to fill the region to inpaint with low rank matrix completion. To this end, a new matrix completion algorithm, better suited to the inpainting application than existing methods, is also developed in this paper. In its simple form, our method does not require any depth prior, unlike most existing light field inpainting algorithms. The method has then been extended to better handle the case where the area to inpaint contains depth discontinuities. In this case, a segmentation map of the different depth layers of the inpainted central view is required. This information is used to warp the depth layers with different homographies. Our experiments with natural light fields captured with plenoptic cameras demonstrate the robustness of the low rank approach to noisy data as well as large color and illumination variations between the views of the light field. Mikael Le Pendu, Xiaoran Jiang, Christine Guillemot |
IEEE Trans. Image Process. | 3 |
| 2017 | JND-Guided Perceptual Pre-filtering for HEVC Compression of UHDTV Video Contents
Eloïse Vidal, François-Xavier Coudoux, Patrick Corlay, Christine Guillemot |
ACIVS | 4 |
| 2017 | Image compression with Stochastic Winner-Take-All Auto-EncoderabstractThis paper addresses the problem of image compression using sparse representations. We propose a variant of autoencoder called Stochastic Winner-Take-All Auto-Encoder (SWTA AE). “Winner-Take-All” means that image patches compete with one another when computing their sparse representation and “Stochastic” indicates that a stochastic hyperparameter rules this competition during training. Unlike auto-encoders, SWTA AE performs variable rate image compression for images of any size after a single training, which is fundamental for compression. For comparison, we also propose a variant of Orthogonal Matching Pursuit (OMP) called Winner-Take-All Orthogonal Matching Pursuit (WTA OMP). In terms of rate-distortion trade-off, SWTA AE outperforms auto-encoders but it is worse than WTA OMP. Besides, SWTA AE can compete with JPEG in terms of rate-distortion. Thierry Dumas, Aline Roumy, Christine Guillemot |
ICASSP | 3 |
| 2017 | Homography-based low rank approximation of light fields for compressionabstractThis paper studies the problem of low rank approximation of light fields for compression. A homography-based approximation method is proposed which jointly searches for homographies to align the different views of the light field together with the low rank approximation matrices. We first consider a global homography per view and show that depending on the variance of the disparity across views, the global homography is not sufficient to well-align the entire images. In a second step, we thus consider multiple homographies, one per region, the region being extracted using depth information. We first show the benefit of the joint optimization of the homographies together with the low-rank approximation. The resulting compact representation is then compressed using HEVC and the results are compared with those obtained by directly applying HEVC on the light field views re-structured as a video sequence. The experiments using different data sets show substantial PSNR-rate gain of our method, especially for real light fields. Xiaoran Jiang, Mikael Le Pendu, Reuben A. Farrugia, Sheila S. Hemami, Christine Guillemot |
ICASSP | 5 |
| 2017 | Scalable light field compression scheme using sparse reconstruction and restorationabstractThis paper describes a light field scalable compression scheme based on the sparsity of the angular Fourier transform of the light field. A subset of sub-aperture images (or views) is compressed using HEVC as a base layer and transmitted to the decoder. An entire light field is reconstructed from this view subset using a method exploiting the sparsity of the light field in the continuous Fourier domain. The reconstructed light field is enhanced using a patch-based restoration method. Then, restored samples are used to predict original ones, in a SHVC-based SNR-scalable scheme. Experiments with different datasets show a significant bit rate reduction of up to 24% in favor of the proposed compression method compared with a direct encoding of all the views with HEVC. The impact of the compression on the quality of the all-in-focus images is also analyzed showing the advantage of the proposed scheme. Fatma Hawary, Christine Guillemot, Dominique Thoreau, Guillaume Boisson |
ICIP | 2 |
| 2017 | Graph-based light fields representation and coding using geometry informationabstractThis paper describes a graph-based coding scheme for light fields (LF). It first adapts graph-based representations (GBR) to describe color and geometry information of LF. Graph connections describing scene geometry capture inter-view dependencies. They are used as the support of a weighted Graph Fourier Transform (wGFT) to encode disoccluded pixels. The quality of the LF reconstructed from the graph is enhanced by adding extra color information to the representation for a sub-set of sub-aperture images. Experiments show that the proposed scheme yields rate-distortion gains compared with HEVC based compression (directly compressing the LF as a video sequence by HEVC). Xin Su 0003, Mira Rizkallah, Thomas Maugey, Christine Guillemot |
ICIP | 4 |
| 2017 | White lenslet image guided demosaicing for plenoptic camerasabstractMost modern cameras use a color filter array on their sensor in order to capture color images. This array is composed of red, green and blue filters and so, each pixel on the sensor lacks two color channels which can be retrieved by a process called demosaicing. In this paper, we propose a new demosaicing method for plenoptic cameras. This type of cameras has become a growing trend and their captured raw images have a particular lenslet structure which must be taken into account to retrieve the sub-aperture images which compose the light field. First, we analyze and describe the flaws of the state-of-the-art light field decoding pipeline. To better identify the different sources of artifacts, our analysis is performed by generating ideal lenslet images from synthetic light fields and use them as input of the decoding pipeline. Then, we detail a new method of demosaicing based on the provided white lenslet images serving as guide. Furthermore, we show that this kind of guided interpolation can be useful on other steps of the decoding pipeline. Finally, the quality of the resulting sub-aperture images is assessed for both synthetic and real light fields using visual comparisons as well as objective metrics. Pierre David 0001, Mikael Le Pendu, Christine Guillemot |
MMSP | 3 |
| 2017 | A Study of the Classification of Low-Dimensional Data with Supervised Manifold Learning
Elif Vural, Christine Guillemot |
J. Mach. Learn. Res. | 2 |
| 2017 | New adaptive filters as perceptual preprocessing for rate-quality performance optimization of video coding
Eloïse Vidal, Nicolas Sturmel, Christine Guillemot, Patrick Corlay, François-Xavier Coudoux |
Signal Process. Image Commun. | 3 |
| 2017 | Scalable Image Coding Based on EpitomesabstractIn this paper, we propose a novel scheme for scalable image coding based on the concept of epitome. An epitome can be seen as a factorized representation of an image. Focusing on spatial scalability, the enhancement layer of the proposed scheme contains only the epitome of the input image. The pixels of the enhancement layer not contained in the epitome are then restored using two approaches inspired from local learning-based super-resolution methods. In the first method, a locally linear embedding model is learned on base layer patches and then applied to the corresponding epitome patches to reconstruct the enhancement layer. The second approach learns linear mappings between pairs of co-located base layer and epitome patches. Experiments have shown that the significant improvement of the rate-distortion performances can be achieved compared with the Scalable extension of HEVC (SHVC). Martin Alain, Christine Guillemot, Dominique Thoreau, Philippe Guillotel |
IEEE Trans. Image Process. | 2 |
| 2017 | Face Hallucination Using Linear Models of Coupled Sparse SupportabstractMost face super-resolution methods assume that low- and high-resolution manifolds have similar local geometrical structure; hence, learn local models on the low-resolution manifold (e.g., sparse or locally linear embedding models), which are then applied on the high-resolution manifold. However, the low-resolution manifold is distorted by the one-to-many relationship between low- and high-resolution patches. This paper presents the Linear Model of Coupled Sparse Support (LM-CSS) method, which learns linear models based on the local geometrical structure on the high-resolution manifold rather than on the low-resolution manifold. For this, in a first step, the low-resolution patch is used to derive a globally optimal estimate of the high-resolution patch. The approximated solution is shown to be close in the Euclidean space to the ground truth, but is generally smooth and lacks the texture details needed by the state-of-the-art face recognizers. Unlike existing methods, the sparse support that best estimates the first approximated solution is found on the high-resolution manifold. The derived support is then used to extract the atoms from the coupled low- and high-resolution dictionaries that are most suitable to learn an up-scaling function for every facial region. The proposed solution was also extended to compute face super-resolution of non-frontal images. Extensive experimental results conducted on a total of 1830 facial images show that the proposed method outperforms seven face super-resolution and a state-of-the-art cross-resolution face recognition method in terms of both quality and recognition. Reuben A. Farrugia, Christine Guillemot |
IEEE Trans. Image Process. | 2 |
| 2017 | Gradient-Based Tone Mapping for Rate-Distortion Optimized Backward-Compatible High Dynamic Range CompressionabstractThis paper addresses the problem of designing a global tone mapping operator for rate distortion optimized backward compatible compression of high dynamic range (HDR) images. We address the problem of tone mapping design for two different use cases leading to two different minimization problems. The first problem considered is the minimization of the distortion on the reconstructed HDR signal under a rate constraint on the standard dynamic range (SDR) layer. The second problem remains the same minimization with an additional constraint to preserve a good quality for the SDR signal. Both the distortion and rate are expressed as a function of the spatial gradient in HDR images. Experiments show that the proposed rate and distortion models based on the HDR image gradient accurately predict the real image rate and distortion measures. Experimental results show that for the first minimization, the optimal rate-distortion performances are achieved, and that the second optimization yields the best tradeoff between rate-distortion performance and quality preservation of the SDR signal. David Gommelet, Aline Roumy, Christine Guillemot, Michaël Ropert, Julien Le Tanou |
IEEE Trans. Image Process. | 3 |
| 2017 | Rate-Distortion Optimized Graph-Based Representation for Multiview Images With Complex Camera ConfigurationsabstractGraph-based representation (GBR) has recently been proposed for describing color and geometry of multiview video content. The graph vertices represent the color information, while the edges represent the geometry information, i.e., the disparity, by connecting corresponding pixels in two camera views. In this paper, we generalize the GBR to multiview images with complex camera configurations. Compared with the existing GBR, the proposed representation can handle not only horizontal displacements of the cameras but also forward/backward translations, rotations, etc. However, contrary to the usual disparity that is a 2-D vector (denoting horizontal and vertical displacements), each edge in GBR is represented by a 1-D disparity. This quantity can be seen as the disparity along an epipolar segment. In order to have a sparse (i.e., easy to code) graph structure, we propose a rate-distortion model to select the most meaningful edges. Hence the graph is constructed with "just enough" information for rendering the given predicted view. The experiments show that the proposed GBR allows high reconstruction quality with lower or equivalent coding rate than traditional depth-based representations. Xin Su 0003, Thomas Maugey, Christine Guillemot |
IEEE Trans. Image Process. | 3 |
| 2016 | Light Field Segmentation Using a Ray-Based Graph Structure
Matthieu Hog, Neus Sabater, Christine Guillemot |
ECCV (7) | 3 |
| 2016 | Learning clustering-based linear mappings for quantization noise removalabstractThis paper describes a novel scheme to reduce the quantization noise of compressed videos and improve the overall coding performances. The proposed scheme first consists in clustering noisy patches of the compressed sequence. Then, at the encoder side, linear mappings are learned for each cluster between the noisy patches and the corresponding source patches. The linear mappings are then transmitted to the decoder where they can be applied to perform de-noising. The method has been tested with the HEVC standard, leading to a bitrate saving of up to 9.63%. Martin Alain, Christine Guillemot, Dominique Thoreau, Philippe Guillotel |
ICIP | 2 |
| 2016 | Robust face hallucination using quantization-adaptive dictionariesabstractExisting face hallucination methods are optimized to super-resolve uncompressed images and are not able to handle the distortions caused by compression. This work presents a new dictionary construction method which jointly models both distortions caused by down-sampling and compression. The resulting dictionaries are then used to make three face super-resolution methods more robust to compression. Experimental results show that the proposed dictionary construction method generates dictionaries which are more representative of the low-quality face image being restored and makes the extended face hallucination methods more robust to compression. These experiments demonstrate that the proposed robust face hallucination methods can achieve Peak Signal-to-Noise Ratio (PSNR) gains between 2-4.48dB and recognition improvement between 2.9-8.1% compared with the low-quality image and outperforming traditional super-resolution methods in most cases. Reuben A. Farrugia, Christine Guillemot |
ICIP | 2 |
| 2016 | Rate-distortion optimization of a tone mapping with SDR quality constraint for backward-compatible high dynamic range compressionabstractThis paper addresses the problem of designing a global tone mapping operator for rate-distortion optimized backward compatible compression of HDR images. We consider a two layer coding scheme in which a base SDR layer is coded with HEVC, inverse tone mapped and subtracted from the input HDR signal to yield the enhancement HDR layer. The tone mapping curve design is formulated as the minimization of the distortion on the reconstructed HDR signal under the constraint of a total rate cost on both layers, while preserving a good quality for the SDR signal. We first demonstrate that the optimum tone mapping function only depends on the rate of the base SDR layer and that the minimization problem can be separated in two consecutive minimization steps. Experimental results show that the proposed tone mapping optimization yields the best trade-off between rate-distortion performance and quality preservation of the coded SDR. David Gommelet, Aline Roumy, Christine Guillemot, Michaël Ropert, Julien Le Tanou |
ICIP | 3 |
| 2016 | Graph-based representation for multiview images with complex camera configurationsabstractInstead of lossily coding depth images resulting in undesirable geometric distortion, graph-based representation (GBR) describes disparity information as a graph with a controllable accuracy. In this paper, we propose a more compact graphical representation called GBR-plus to code both disparity and color information of a target view given a reference view. Specifically, first we differentiate between disocclusion holes (occluded spatial regions in the reference view) and rounding holes (insufficiently sampled regions in the reference view) in the synthesized target view, so that the decoder can optionally complete rounding holes via signal interpolation without coding overhead. Second, we use a compact graphical representation to delimit disparity-shifted boundaries of objects in the target view, which is coded losslessly. Finally, color pixels in disocclusion holes are predicted using adjacent background pixels as predictors, and prediction residuals in a local neighborhood are coded using Graph Fourier Transform (GFT). Experimental results show that GBR-plus outperforms previous GBR, and has comparable performance as HEVC at mid to high bitrates with lower encoder complexity. Xin Su 0003, Thomas Maugey, Christine Guillemot |
ICIP | 3 |
| 2016 | Geometry-Aware Neighborhood Search for Learning Local Models for Image SuperresolutionabstractLocal learning of sparse image models has proved to be very effective to solve inverse problems in many computer vision applications. To learn such models, the data samples are often clustered using the K-means algorithm with the Euclidean distance as a dissimilarity metric. However, the Euclidean distance may not always be a good dissimilarity measure for comparing data samples lying on a manifold. In this paper, we propose two algorithms for determining a local subset of training samples from which a good local model can be computed for reconstructing a given input test sample, where we consider the underlying geometry of the data. The first algorithm, called adaptive geometry-driven nearest neighbor search (AGNN), is an adaptive scheme, which can be seen as an out-of-sample extension of the replicator graph clustering method for local model learning. The second method, called geometry-driven overlapping clusters (GOCs), is a less complex nonadaptive alternative for training subset selection. The proposed AGNN and GOC methods are evaluated in image superresolution and shown to outperform spectral clustering, soft clustering, and geodesic distance-based subset selection in most settings. Júlio César Ferreira, Elif Vural, Christine Guillemot |
IEEE Trans. Image Process. | 3 |
| 2016 | Inter-Layer Prediction of Color in High Dynamic Range Image Scalable CompressionabstractThis paper presents a color inter-layer prediction (ILP) method for scalable coding of high dynamic range (HDR) video content with a low dynamic range (LDR) base layer. Relying on the assumption of hue preservation between the colors of an HDR image and its LDR tone mapped version, we derived equations for predicting the chromatic components of the HDR layer given the decoded LDR layer. Two color representations are studied. In a first encoding scheme, the HDR image is represented in the classical Y'CbCr format. In addition, a second scheme is proposed using a colorspace based on the CIE u'v' uniform chromaticity scale diagram. In each case, different prediction equations are derived based on a color model ensuring the hue preservation. Our experiments highlight several advantages of using a CIE u'v'-based colorspace for the compression of HDR content, especially in a scalable context. In addition, our ILP scheme using this color representation improves on the state-of-the-art ILP method, which directly predicts the HDR layer u'v' components by computing the LDR layers u'v' values of each pixel. Mikael Le Pendu, Christine Guillemot, Dominique Thoreau |
IEEE Trans. Image Process. | 2 |
| 2016 | Out-of-Sample Generalizations for Supervised Manifold Learning for ClassificationabstractSupervised manifold learning methods for data classification map high-dimensional data samples to a lower dimensional domain in a structure-preserving way while increasing the separation between different classes. Most manifold learning methods compute the embedding only of the initially available data; however, the generalization of the embedding to novel points, i.e., the out-of-sample extension problem, becomes especially important in classification applications. In this paper, we propose a semi-supervised method for building an interpolation function that provides an out-of-sample extension for general supervised manifold learning algorithms studied in the context of classification. The proposed algorithm computes a radial basis function interpolator that minimizes an objective function consisting of the total embedding error of unlabeled test samples, defined as their distance to the embeddings of the manifolds of their own class, as well as a regularization term that controls the smoothness of the interpolation function in a direction-dependent way. The class labels of test data and the interpolation function parameters are estimated jointly with an iterative process. Experimental results on face and object images demonstrate the potential of the proposed out-of-sample extension algorithm for the classification of manifold-modeled data sets. Elif Vural, Christine Guillemot |
IEEE Trans. Image Process. | 2 |
| 2015 | Mode Dependent Vector Quantization with a rate-distortion optimized codebook for residue coding in video compressionabstractThe High Efficiency Video Coding standard (HEVC) supports a total of 35 intra prediction modes which aim at reducing spatial redundancy by exploiting pixel correlation within a local neighborhood. In this paper, we show that spatial correlation remains after intra prediction, leading to high energy prediction residues. We propose a novel scheme for encoding the prediction residues using a Mode Dependent Vector Quantization (MDVQ) which aims at reducing the redundancy in residual domain. The MDVQ codebook is optimized in a rate-distortion (RD) sense. Experimental results show that the codebook can be independent of the quantization parameter (QP) with no loss in terms of coding efficiency. A bitrate reduction of 1.1% on average compared to HEVC can be achieved, while further tests indicate that codebook adaptivity could substantially improve the performance. Bihong Huang, Félix Henry, Christine Guillemot, Philippe Salembier |
ICASSP | 3 |
| 2015 | Guided inpainting with cluster-based auxiliary informationabstractIn this paper, we propose a new guided inpainting algorithm based on the exemplar-based approach in order to effectively fill in holes in image synthesis applications. Guided inpainting techniques can be very useful in settings where one has access to the ground truth information like most multiview coding applications. We propose a new auxiliary information based on patch clustering, which is used to refine the candidate exemplar set in the inpainting. For that purpose, a new recursive clustering method based on locally linear embedding (LLE) is introduced. We then design the guided inpainting solution based on LLE with clustered patches, which contrains the reconstruction to operate in one patch cluster only. The index of the appropriate cluster considered as auxiliary information. Experimental results show that our clustering algorithm provides clusters that are well suited to the inpainting problem. They also show that the auxiliary information enables to significantly improve the quality of the inpainted image for a small coding cost. This work is the first study to show that effective inpainting can be performed when the auxiliary information is properly adapted to the characteristics of both the hole and the known texture. Thomas Maugey, Pascal Frossard, Christine Guillemot |
ICIP | 3 |
| 2015 | Template based inter-layer prediction for high dynamic range scalable compressionabstractThis paper presents a scalable high dynamic range (HDR) image coding framework in which the base layer is a low dynamic range (LDR) version of the image that may have been generated by an arbitrary Tone Mapping Operator (TMO). Our method successfully handles the case of complex local TMOs thanks to a block-wise and non-linear approach. A novel template based Inter Layer Prediction (ILP) is designed in order to perform the inverse tone mapping of a block without the need to transmit any additional parameter to the decoder. This method enables the use of a more accurate inverse tone mapping model than the simple linear regression commonly used for block-wise ILP. Our experiments have shown an average bitrate saving of 34% on the HDR enhancement layer, compared to state of the art methods. Mikael Le Pendu, Christine Guillemot, Dominique Thoreau |
ICIP | 2 |
| 2015 | Epitomic image factorization via neighbor-embeddingabstractWe describe a novel epitomic image representation scheme that factors a given image content into a condensed epitome and a low-resolution image to reduce the memory space for images. Given an input image, we construct a condensed epitome such that all image patches can successfully be reconstructed from the factored representation by means of an optimized neighbor-embedding strategy. Under this new scope of epitomic image representations aligned with the manifold sampling assumption, we end up a more generic epitome learning scheme with increased optimality, compactness, and reconstruction stability. We present the performance of the proposed method for image and video up-scaling (super-resolution) while extensions to other image and video processing are straightforward. Mehmet Türkan, Martin Alain, Dominique Thoreau, Philippe Guillotel, Christine Guillemot |
ICIP | 5 |
| 2015 | Inter-prediction methods based on linear embedding for video compression
Martin Alain, Christine Guillemot, Dominique Thoreau, Philippe Guillotel |
Signal Process. Image Commun. | 2 |
| 2015 | Video Inpainting With Short-Term Windows: Application to Object Removal and Error ConcealmentabstractIn this paper, we propose a new video inpainting method which applies to both static or free-moving camera videos. The method can be used for object removal, error concealment, and background reconstruction applications. To limit the computational time, a frame is inpainted by considering a small number of neighboring pictures which are grouped into a group of pictures (GoP). More specifically, to inpaint a frame, the method starts by aligning all the frames of the GoP. This is achieved by a region-based homography computation method which allows us to strengthen the spatial consistency of aligned frames. Then, from the stack of aligned frames, an energy function based on both spatial and temporal coherency terms is globally minimized. This energy function is efficient enough to provide high quality results even when the number of pictures in the GoP is rather small, e.g. 20 neighboring frames. This drastically reduces the algorithm complexity and makes the approach well suited for near real-time video editing applications as well as for loss concealment applications. Experiments with several challenging video sequences show that the proposed method provides visually pleasing results for object removal, error concealment, and background reconstruction context. Mounira Ebdelli, Olivier Le Meur, Christine Guillemot |
IEEE Trans. Image Process. | 3 |
| 2015 | Local Inverse Tone Curve Learning for High Dynamic Range Image Scalable CompressionabstractThis paper presents a scalable high dynamic range (HDR) image coding scheme in which the base layer is a low dynamic range version of the image that may have been generated by an arbitrary tone mapping operator (TMO). No restriction is imposed on the TMO, which can be either global or local, so as to fully respect the artistic intent of the producer. Our method successfully handles the case of complex local TMOs thanks to a block-wise and non-linear approach. A novel template-based interlayer prediction (ILP) is designed in order to perform the inverse tone mapping of a block without the need to transmit any additional parameter to the decoder. This method enables the use of a more accurate inverse tone mapping model than the simple linear regression commonly used for block-wise ILP. In addition, this paper shows that a linear adjustment of the initially predicted block can further improve the overall coding performance by using an efficient encoding scheme of the scaling parameters. Our experiments have shown an average bitrate saving of 47% on the HDR enhancement layer, compared with the previous local ILP methods. Mikael Le Pendu, Christine Guillemot, Dominique Thoreau |
IEEE Trans. Image Process. | 2 |
| 2014 | Adaptive re-quantization for high dynamic range video compressionabstractHigh Dynamic Range (HDR) images contain more intensity levels than traditional image formats. Instead of 8 or 10 bit integers, floating point values are generally used to represent the pixel data. To extend the use of existing video codecs such as HEVC to HDR floating point video sequences, we propose a method that converts the floating point data and reduces the bit depth of input images with minimal loss. Several variants of the method are proposed. They are adapted to different quality requirements. In particular, near lossless compression is addressed. Mikael Le Pendu, Christine Guillemot, Dominique Thoreau |
ICASSP | 2 |
| 2014 | Epitome inpainting with in-loop residue coding for image compressionabstractThis paper describes a novel image coding scheme based on epitome inpainting. An epitome containing a factorized texture representation of the image is first coded and transmitted. The decoded epitome is then inpainted by propagating its structure and texture with extended H.264 Intra directional prediction modes and advanced neighbor embedding methods. The blocks to be inpainted are processed according to an adaptive scanning order which adapts to the image content and gives a higher priority to the blocks having strong structures. The inpainting step has been incorporated in the coding loop of an H.264 encoder, leading to a bitrate saving of up to 48.12% with respect to H.264 intra. Safa Chérigui, Martin Alain, Christine Guillemot, Dominique Thoreau, Philippe Guillotel |
ICIP | 3 |
| 2014 | HEVC Intra coding of ultra HD video with reduced complexityabstractThe HEVC (High Efficiency Video Coding) standard brings the necessary quality versus rate performance for efficient transmission of Ultra High Definition formats (UHD). However, one of the remaining barriers to its adoption for UHD content is its high encoding complexity. In this paper, we address the problem of HEVC encoding complexity reduction by proposing a strategy to infer UHD coding modes and quadtree structure from those optimized for the lower (HD) resolution version of the input video. A speed-up factor of 3 is achieved compared to directly encoding the UHD format at the expense of a limited quality loss. Nicolas Dhollande, Olivier Le Meur, Christine Guillemot |
ICIP | 3 |
| 2014 | Single image super-resolution using sparse representations with structure constraintsabstractThis paper describes a new single-image super-resolution algorithm based on sparse representations with image structure constraints. A structure tensor based regularization is introduced in the sparse approximation in order to improve the sharpness of edges. The new formulation allows reducing the ringing artefacts which can be observed around edges reconstructed by existing methods. The proposed method, named Sharper Edges based Adaptive Sparse Domain Selection (SE-ASDS), achieves much better results than many state-of-the-art algorithms, showing significant improvements in terms of PSNR (average of 29.63, previously 29.19), SSIM (average of 0.8559, previously 0.8471) and visual quality perception. Júlio César Ferreira, Olivier Le Meur, Christine Guillemot, Eduardo A. B. da Silva, Gilberto Arantes Carrijo |
ICIP | 3 |
| 2014 | Inter-view motion prediction in 3D-HEVCabstractThis paper presents a novel inter-view motion prediction technique used in 3D-HEVC which provides efficient compression for motion vectors. The 3D extension of HEVC (High Efficiency Video Coding) namely 3D-HEVC is under development in JCT-3V for coding multi-view video and depth data. Inter-view motion prediction takes benefit of the inter-view correlation between views by inferring the motion information of a view from the already coded motion information in another view as well as the local disparity between two views. Experimental results show that inter-view motion prediction in 3D-HEVC provides an average bit-rate savings of 15% for the dependent views. Li Zhang 0006, Ying Chen 0011, Vijayaraghavan Thirumalai, Jian-Liang Lin, Yi-Wen Chen, Jicheng An, Shawmin Lei, Laurent Guillo, Thomas Guionnet, Christine Guillemot |
ISCAS | 10 |
| 2014 | Single-Image Super-Resolution via Linear Mapping of Interpolated Self-ExamplesabstractThis paper presents a novel example-based single-image superresolution procedure that upscales to high-resolution (HR) a given low-resolution (LR) input image without relying on an external dictionary of image examples. The dictionary instead is built from the LR input image itself, by generating a double pyramid of recursively scaled, and subsequently interpolated, images, from which self-examples are extracted. The upscaling procedure is multipass, i.e., the output image is constructed by means of gradual increases, and consists in learning special linear mapping functions on this double pyramid, as many as the number of patches in the current image to upscale. More precisely, for each LR patch, similar self-examples are found, and, because of them, a linear function is learned to directly map it into its HR version. Iterative back projection is also employed to ensure consistency at each pass of the procedure. Extensive experiments and comparisons with other state-of-the-art methods, based both on external and internal dictionaries, show that our algorithm can produce visually pleasant upscalings, with sharp edges and well reconstructed details. Moreover, when considering objective metrics, such as Peak signal-to-noise ratio and Structural similarity, our method turns out to give the best performance. Marco Bevilacqua, Aline Roumy, Christine Guillemot, Marie-Line Alberi-Morel |
IEEE Trans. Image Process. | 3 |
| 2013 | Compact and coherent dictionary construction for example-based super-resolutionabstractThis paper presents a new method to construct a dictionary for example-based super-resolution (SR) algorithms. Example-based SR relies on a dictionary of correspondences of low-resolution (LR) and high-resolution (HR) patches. Having a fixed, prebuilt, dictionary, allows to speed up the SR process; however, in order to perform well in most cases, we need to have big dictionaries with a large variety of patches. Moreover, LR and HR patches often are not coherent, i.e. local LR neighborhoods are not preserved in the HR space. Our designed dictionary learning method takes as input a large dictionary and gives as an output a dictionary with a “sustainable” size, yet presenting comparable or even better performance. It firstly consists of a partitioning process, done according to a joint k-means procedure, which enforces the coherence between LR and HR patches by discarding those pairs for which we do not find a common cluster. Secondly, the clustered dictionary is used to extract some salient patches that will form the output set. Marco Bevilacqua, Aline Roumy, Christine Guillemot, Marie-Line Alberi-Morel |
ICASSP | 3 |
| 2013 | Analysis of patch-based similarity metrics: Application to denoisingabstractThis paper presents a performance analysis of measures used for assessing similarities between patches. Compared to subjective ground thruth, our results indicate that some metrics are more suitable than others in a context of patch matching. This conclusion is confirmed by an experiment on non-local means (NLM) denoising algorithm. The denoising quality depends on the chosen similarity metric. In the best case, the gain is of 1.3dB compared to a classical SSD-based denoising algorithm. Mounira Ebdelli, Olivier Le Meur, Christine Guillemot |
ICASSP | 3 |
| 2013 | K-NN search using local learning based on regression for neighbor embedding-based image predictionabstractThe paper describes a K-NN search method aided by local learning of subspace mappings for the problem of neighbor-embedding based image Intra prediction. The local learning of subspace mappings relies on multivariate linear regression. The method is used jointly with Locally Linear Embedding (LLE) as well as with a method inspired from Non Local Means (NLM) for prediction. Linear and kernel ridge regression are also considered directly for predicting the unknown pixels. Rate-distortion performances are then given in comparison with Intra prediction using LLE and classical K-NN search, as well as in comparison with H.264 Intra prediction modes. Christine Guillemot, Safa Chérigui, Dominique Thoreau |
ICASSP | 1 |
| 2013 | Image inpainting using LLE-LDNR and linear subspace mappingsabstractThe paper first describes an examplar-based image inpainting algorithm using a locally linear neighbor embedding technique with low-dimensional neighborhood representation (LLE-LDNR). The inpainting algorithm first searches the K nearest neighbors ( ) of the input patch to be filled-in and linearly combine them with LLE-LDNR to synthesize the missing pixels. Linear regression is then introduced for improving the K-NN search. The performance of the LLE-LDNR with the enhanced K-NN search method is assessed for two applications: loss concealment and object removal. Christine Guillemot, Mehmet Türkan, Olivier Le Meur, Mounira Ebdelli |
ICASSP | 1 |
| 2013 | Learning a tree-structured dictionary for efficient image representation with adaptive sparse codingabstractWe introduce a new method, called Tree K-SVD, to learn a tree-structured dictionary for sparse representations, as well as a new adaptive sparse coding method, in a context of image compression. Each dictionary at a level in the tree is learned from residuals from the previous level with the K-SVD method. The tree-structured dictionary allows efficient search of the atoms along the tree as well as efficient coding of their indices. Besides, it is scalable in the sense that it can be used, once learned, for several sparsity constraints. We show experimentally on face images that, for a high sparsity, Tree K-SVD offers better rate-distortion performances than state-of-the-art “flat” dictionaries learned by K-SVD or Sparse K-SVD, or than the predetermined overcomplete DCT dictionary. We also show that our adaptive sparse coding method, used on a tree-structured dictionary to adapt the sparsity per level, improves the quality of reconstruction. Jeremy Aghaei Mazaheri, Christine Guillemot, Claude Labit |
ICASSP | 2 |
| 2013 | Locally linear embedding methods for inter image codingabstractImage prediction methods based on data dimensionality reduction techniques have been recently introduced for still images. These techniques have been proven efficient, especially when used in a H.264 framework. This paper introduces a natural extension of this work from spatial prediction to temporal prediction. Locally Linear Embedding and Optimized Map-Aided Locally Linear Embedding methods are adapted within H.264 in order to improve inter image prediction. The resulting prediction methods are compared with the H.264 motion estimation/compensation method, template matching methods and Adaptive Interpolation Filters, and are shown to bring significant Rate-Distortion performance improvements (up to 20.21 % rate saving with respect to H.264). Martin Alain, Safa Chérigui, Christine Guillemot, Dominique Thoreau, Philippe Guillotel |
ICIP | 3 |
| 2013 | K-WEB: Nonnegative dictionary learning for sparse image representationsabstractThis paper presents a new nonnegative dictionary learning method, to decompose an input data matrix into a dictionary of nonnegative atoms, and a representation matrix with a strict ℓ0-sparsity constraint. This constraint makes each input vector representable by a limited combination of atoms. The proposed method consists of two steps which are alternatively iterated: a sparse coding and a dictionary update stage. As for the dictionary update, an original method is proposed, which we call K-WEB, as it involves the computation of k WEighted Barycenters. The so designed algorithm is shown to outperform other methods in the literature that address the same learning problem, in different applications, and both with synthetic and “real” data, i.e. coming from natural images. Marco Bevilacqua, Aline Roumy, Christine Guillemot, Marie-Line Alberi-Morel |
ICIP | 3 |
| 2013 | Video super-resolution via sparse combinations of key-frame patches in a compression contextabstractIn this paper we present a super-resolution (SR) method for upscaling low-resolution (LR) video sequences, that relies on the presence of periodic high-resolution (HR) key frames, and validate it in the context of video compression. For a given LR intermediate frame, the HR details are retrieved patch-by-patch by taking sparse linear combinations of patches found in the neighbor key frames. The performance of the video SR algorithm is assessed in a scheme where only some key frames from an original HR sequence are directly encoded; the remaining intermediate frames are down-sampled to LR and encoded as well, with a possibly different quantization parameter. SR is then finally employed to upscale these frames. For comparison, we consider the best case where the whole original HR sequence is encoded. With respect to this case, our SR-based approach is shown to bring a certain gain for low bit-rates (consistent when all frames are encoded independently), i.e. when a poor encoding can actually benefit of the special processing of the intermediate frames, so proving that video SR can be an useful tool in realistic scenarios. Marco Bevilacqua, Aline Roumy, Christine Guillemot, Marie-Line Alberi-Morel |
PCS | 3 |
| 2013 | Learning an adaptive dictionary structure for efficient image sparse codingabstractWe introduce a new method to learn an adaptive dictionary structure suitable for efficient coding of sparse representations. The method is validated in a context of satellite image compression. The dictionary structure adapts itself during the learning to the training data and can be seen as a tree structure whose branches are progressively pruned depending on their usage rate and merged into a single branch. This adaptive structure allows a fast search for the atoms and an efficient coding of their indices. It is also scalable in sparsity, meaning that once learned, the structure can be used for several sparsity values. We show experimentally that this adaptive structure offers better rate-distortion performances than the “flat” K-SVD dictionary, a dictionary structured in one branch, and the tree-structured K-SVD dictionary (called Tree K-SVD). Jeremy Aghaei Mazaheri, Christine Guillemot, Claude Labit |
PCS | 2 |
| 2013 | Dictionary learning for image prediction
Mehmet Türkan, Christine Guillemot |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Object removal and loss concealment using neighbor embedding methods
Christine Guillemot, Mehmet Türkan, Olivier Le Meur, Mounira Ebdelli |
Signal Process. Image Commun. | 1 |
| 2013 | Correspondence Map-Aided Neighbor Embedding for Image Intra PredictionabstractThis paper describes new image prediction methods based on neighbor embedding (NE) techniques. Neighbor embedding methods are used here to approximate an input block (the block to be predicted) in the image as a linear combination of K nearest neighbors. However, in order for the decoder to proceed similarly, the K nearest neighbors are found by computing distances between the known pixels in a causal neighborhood (called template) of the input block and the co-located pixels in candidate patches taken from a causal window. Similarly, the weights used for the linear approximation are computed in order to best approximate the template pixels. Although efficient, these methods suffer from limitations when the template and the block to be predicted are not correlated, e.g., in non homogenous texture areas. To cope with these limitations, this paper introduces new image prediction methods based on NE techniques in which the K-NN search is done in two steps and aided, at the decoder, by a block correspondence map, hence the name map-aided neighbor embedding (MANE) method. Another optimized variant of this approach, called oMANE method, is also studied. In these methods, several alternatives have also been proposed for the K-NN search. The resulting prediction methods are shown to bring significant rate-distortion performance improvements when compared to H.264 Intra prediction modes (up to 44.75% rate saving at low bit rates). Safa Chérigui, Christine Guillemot, Dominique Thoreau, Philippe Guillotel, Patrick Pérez |
IEEE Trans. Image Process. | 2 |
| 2013 | Hierarchical Super-Resolution-Based InpaintingabstractThis paper introduces a novel framework for examplar-based inpainting. It consists in performing first the inpainting on a coarse version of the input image. A hierarchical super-resolution algorithm is then used to recover details on the missing areas. The advantage of this approach is that it is easier to inpaint low-resolution pictures than high-resolution ones. The gain is both in terms of computational complexity and visual quality. However, to be less sensitive to the parameter setting of the inpainting method, the low-resolution input picture is inpainted several times with different configurations. Results are efficiently combined with a loopy belief propagation and details are recovered by a single-image super-resolution algorithm. Experimental results in a context of image editing and texture synthesis demonstrate the effectiveness of the proposed method. Results are compared to five state-of-the-art inpainting methods. Olivier Le Meur, Mounira Ebdelli, Christine Guillemot |
IEEE Trans. Image Process. | 3 |
| 2012 | Low-Complexity Single-Image Super-Resolution based on Nonnegative Neighbor EmbeddingabstractInternational audience Marco Bevilacqua, Aline Roumy, Christine Guillemot, Marie-Line Alberi-Morel |
BMVC | 3 |
| 2012 | Super-Resolution-Based Inpainting
Olivier Le Meur, Christine Guillemot |
ECCV (6) | 2 |
| 2012 | Neighbor embedding based single-image super-resolution using Semi-Nonnegative Matrix FactorizationabstractThis paper describes a novel method for single-image super-resolution (SR) based on a neighbor embedding technique which uses Semi-Nonnegative Matrix Factorization (SNMF). Each low-resolution (LR) input patch is approximated by a linear combination of nearest neighbors taken from a dictionary. This dictionary stores low-resolution and corresponding high-resolution (HR) patches taken from natural images and is thus used to infer the HR details of the super-resolved image. The entire neighbor embedding procedure is carried out in a feature space. Features which are either the gradient values of the pixels or the mean-subtracted luminance values are extracted from the LR input patches, and from the LR and HR patches stored in the dictionary. The algorithm thus searches for the K nearest neighbors of the feature vector of the LR input patch and then computes the weights for approximating the input feature vector. The use of SNMF for computing the weights of the linear approximation is shown to have a more stable behavior than the use of LLE and lead to significantly higher PSNR values for the super-resolved images. Marco Bevilacqua, Aline Roumy, Christine Guillemot, Marie-Line Alberi-Morel |
ICASSP | 3 |
| 2012 | Hybrid template and block matching algorithm for image intra predictionabstractTemplate matching has been shown to outperform the H.264 prediction modes for Intra video coding thanks to better spatial prediction and no additional ancillary data to transmit. The method indeed works well when the template and the block to be predicted are highly correlated, e.g., in homogenous image areas, however, it obviously fails in areas with non homogeneous textures. This paper explores the idea of using a block-matching intra prediction algorithm which, thanks to a Rate-Distorsion (RD) based decision mechanism, will naturally be used in image areas when template matching (TM) fails. This new method offers a significant coding gain compared to H.264 Intra prediction modes and the template matching based prediction. Indeed, the TM-based algorithm and the proposed hybrid algorithm lead, with the Bjontergaard measure, to rate gains of up to respectively 38.02% and 48.38% at low bitrates when compared with H.264 Intra only. Safa Chérigui, Christine Guillemot, Dominique Thoreau, Philippe Guillotel, Patrick Pérez |
ICASSP | 2 |
| 2012 | Neighbor embedding with non-negative matrix factorization for image predictionabstractThe paper studies several non-negative matrix factorization methods with nearest neighbors constrained dictionaries for image prediction. The methods considered include the multiplicative update algorithm, the projected gradient algorithm, as well as the graph-regularized NMF solution which aims at taking into account the geometrical structure of the input data. The Intra prediction problem based on these NMF solutions amounts to a neighbor embedding problem. Both prediction and rate-distortion performances are then given in comparison with other neighbor embedding methods like locally linear embedding (LLE) and locally linear embedding with low dimensional neigborhood representation (LLE-LDNR). Christine Guillemot, Mehmet Türkan |
ICASSP | 1 |
| 2012 | Map-Aided Locally Linear Embedding methods for image predictionabstractImage prediction methods based on data dimensionality reduction techniques have been introduced in [1]. Although efficient, these methods suffer from limitations when the block to be predicted and its neighborhood (or template) are not correlated, e.g. in non homogenous texture areas. To cope with these limitations, this paper introduces new image prediction methods based on locally linear embedding (LLE) technique in which the required K-NN search is aided, at the decoder, by a block correspondence map, hence the name Map-Aided Locally Linear Embedding (MALLE) method. Another optimized variant of this approach, called oMALLE method, is also studied. The resulting prediction methods are shown to bring significant Rate-Distortion (RD) performance improvements when compared to H.264 Intra prediction modes (up to 40.78 % rate saving at low bit rates). Safa Chérigui, Christine Guillemot, Dominique Thoreau, Philippe Guillotel, Patrick Pérez |
ICIP | 2 |
| 2012 | Examplar-based video inpainting with motion-compensated neighbor embeddingabstractThis paper describes a video inpainting algorithm based on motion-compensated neighbor embedding. The unknown pixels are estimated as a linear combination of the K closest patches using motion-compensated neighbor embedding. The algorithm is first assessed by assuming the motion information of the masked pixels to be known. This assumption is not realistic in video editing (object removal) applications. It however helps isolating the various problems for the sake of analysis. Different approaches are then assessed in the context where the motion information of missing pixels is unknown. Experiments on several videos show the benefits of the proposed approach which lead to natural looking videos with less annoying artefacts than when using a template matching technique. Mounira Ebdelli, Christine Guillemot, Olivier Le Meur |
ICIP | 2 |
| 2012 | Locally linear embedding based texture synthesis for image prediction and error concealmentabstractThe template matching algorithm is a simple extension to exemplar-based texture synthesis. Average of template matching predictors or non-local means based approaches can be seen as heuristic extensions to template matching. These methods which linearly combine several texture patches have been shown to be more robust in synthesis and to give better results when compared to simple template matching. However, they do not search to minimize an approximation error on the known pixel values in the template. They are rather heuristic methods for calculating the linear weighting coefficients. This paper proposes a neighbor embedding based texture synthesis method by formulating the problem as a least-squares optimization using locally linear embedding. By this means, one calculates the linear weighting coefficients by solving a constrained optimization for approximating the template. The proposed texture synthesis framework has first been applied to the image prediction (predictive coding) problem. It has then been extended to a loss concealment application for transmission errors. Experimental results demonstrate the effectiveness of the proposed method for both image compression and error concealment. Mehmet Türkan, Christine Guillemot |
ICIP | 2 |
| 2012 | Efficient depth map compression based on lossless edge coding and diffusionabstractThe multi-view plus depth video (MVD) format has recently been introduced for 3DTV and free-viewpoint video (FVV) scene rendering. Given one view (or several views) with its depth information, depth image-based rendering techniques have the ability to generate intermediate views. The MVD format however generates large volumes of data which need to be compressed for storage and transmission. This paper describes a new depth map encoding algorithm which aims at exploiting the intrinsic depth maps properties. Depth images indeed represent the scene surface and are characterized by areas of smoothly varying grey levels separated by sharp edges at the position of object boundaries. Preserving these characteristics is important to enable high quality view rendering at the receiver side. The proposed algorithm proceeds in three steps: the edges at object boundaries are first detected using a Sobel operator. The positions of the edges are encoded using the JBIG algorithm. The luminance values of the pixels along the edges are then encoded using an optimized path encoder. The decoder runs a fast diffusion-based inpainting algorithm which fills in the unknown pixels within the objects by starting from their boundaries. The performance of the algorithm is assessed against JPEG-2000 and HEVC, both in terms of PSNR of the depth maps versus rate as well as in terms of PSNR of the synthesized virtual views. Josselin Gautier, Olivier Le Meur, Christine Guillemot |
PCS | 3 |
| 2012 | Source Modeling for Distributed Video CodingabstractThis paper studies source and correlation models for distributed video coding (DVC). It first considers a two-state HMM, i.e., a Gilbert-Elliott process, to model the bit-planes produced by DVC schemes. A statistical analysis shows that this model allows us to accurately capture the memory present in the video bit-planes. The achievable rate bounds are derived for these ergodic sources, first assuming an additive binary symmetric correlation channel between the two sources. These bounds show that a rate gain can be achieved by exploiting the sources memory with the additive BSC model. A Slepian-Wolf decoding algorithm which jointly estimates the sources and the source model parameters is then described. Simulation results show that the additive correlation model does not always fit well with the correlation between the actual video bit-planes. This has led us to consider a second correlation model (the predictive model). The rate bounds are then derived for the predictive correlation model in the case of memory sources, showing that exploiting the source memory does not bring any rate gain and that the noise statistic is a sufficient statistic for the MAP decoder. We also evaluate the rate loss when the correlation model assumed by the decoder is not matched to the true one. An a posteriori estimation of the correlation channel has hence been added to the decoder in order to use the most appropriate correlation model for each bit-plane. The new decoding algorithm has been integrated in a DVC decoder, leading to a rate saving of up to 10.14% for the same PSNR, with respect to the case where the bit-planes are assumed to be memoryless uniform sources correlated with the SI via an additive channel model. Velotiaray Toto-Zarasoa, Aline Roumy, Christine Guillemot |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Image Prediction Based on Neighbor-Embedding MethodsabstractThis paper describes two new intraimage prediction methods based on two data dimensionality reduction methods: nonnegative matrix factorization (NMF) and locally linear embedding. These two methods aim at approximating a block to be predicted in the image as a linear combination of k-nearest neighbors determined on the known pixels in a causal neighborhood of the input block. Variable k can be seen as a parameter controlling some sort of sparsity constraints of the approximation vector. The impact of this parameter as well as of the nonnegativity and sum-to-one constraints for the addressed prediction problem has been analyzed. The prediction and RD performances of these two new image prediction methods have then been evaluated in a complete image coding-and-decoding algorithm. Simulation results show gains up to 2 dB in terms of the PSNR of the reconstructed signal after coding and decoding of the prediction residue when compared with H.264/AVC intraprediction modes, up to 3 dB when compared with template matching, and up to 1 dB when compared with a sparse prediction method. Mehmet Türkan, Christine Guillemot |
IEEE Trans. Image Process. | 2 |
| 2011 | Image prediction based on non-negative matrix factorizationabstractThis paper presents a novel spatial texture prediction method based on non-negative matrix factorization. As an extension of template matching, approximation based iterative texture prediction methods have recently been considered for image prediction. These approaches rely on the assumption that the given basis functions (atoms) span the signal residue space at each iteration of the algorithm. However, in the case of signal prediction with a sup port region approximation, the atoms may not approximate residue signals very well even though the dictionary has been well adapted in the spatial domain. The underlying main idea is to consider a factorization based algorithm in which the given atoms approximate the signal without going further into signal residue space. The proposed spatial prediction method has first been assessed against the prediction methods based on template matching and sparse approximations. It has then been assessed in a compression scheme where the prediction residue is transform encoded. Experimental results obtained show that the proposed method outperforms the template matching and sparse approximations based techniques in terms of encoding efficiency. Mehmet Türkan, Christine Guillemot |
ICASSP | 2 |
| 2011 | Image compression using the Iteration-Tuned and Aligned DictionaryabstractWe present a new, block-based image codec based on sparse representations using a learned, structured dictionary called the Iteration Tuned and Aligned Dictionary (ITAD). The question of selecting the number of atoms used in the representation of each image block is addressed with a new, global (image-wide), rate-distortion-based sparsity selection criterion. We show experimentally that our codec outperforms JPEG2000 in both quantitative evaluations (by 0.9 dB to 4 dB) and qualitative evaluations. Joaquin Zepeda, Christine Guillemot, Ewa Kijak |
ICASSP | 2 |
| 2011 | Object-based Layered Depth Images for improved virtual view synthesis in rate-constrained contextabstractLayered Depth Image (LDI) representations are attractive compact representations for multi-view videos. Any virtual viewpoint can be rendered from LDI by using view synthesis technique. However, rendering from classical LDI leads to annoying visual artifacts, such as cracks and disocclusions. Visual quality gets even worse after a DCT-based compression of the LDI, because of blurring effects on depth discontinuities. In this paper, we propose a novel object-based LDI representation, improving synthesized virtual views quality, in a rate-constrained context. Pixels from each LDI layer are reorganised to enhance depth continuity. Vincent Jantet, Christine Guillemot, Luce Morin |
ICIP | 2 |
| 2011 | Examplar-based inpainting based on local geometryabstractIn this paper, we propose a novel inpainting algorithm combining the advantages of PDE-based schemes and examplar-based approaches. The proposed algorithm relies on the use of structure tensors to define the filling order priority and template matching. The structure tensors are computed in a hierarchic manner whereas the template matching is based on a K-nearest neighbor algorithm. The value K is adaptively set in function of the local texture information. Compared to two state of the art approaches, the proposed method provides more coherent results. Olivier Le Meur, Josselin Gautier, Christine Guillemot |
ICIP | 3 |
| 2011 | Online dictionaries for image predictionabstractThis paper presents a novel dictionary learning method which, because of its simplicity and the limited number of training samples it requires, can be used for online learning of dictionaries for spatial texture prediction. The proposed learning method has first been described to address the problem of intra image prediction based on signal expansion on overcomplete dictionaries. It has then been evaluated in a complete image codec. The experimental results obtained show a significant improvement in terms of the quality of the predicted image compared to H.264/AVC intra prediction. Significant rate-distortion gains have also been achieved on the reconstructed image, after coding and decoding the prediction residue, compared with the H.264/AVC and a sparse spatial prediction method which will be referred to as the generalized template matching approach. Mehmet Türkan, Christine Guillemot |
ICIP | 2 |
| 2011 | Epitome-based image compression using translational sub-pel mappingabstractThis paper addresses the problem of epitome construction for image compression. An optimized epitome construction method is first described, where the epitome and the associated image reconstruction, are both successively performed at full pel and sub-pel accuracy. The resulting complete still image compression scheme is then discussed with details on some innovative tools. The PSNR-rate performance achieved with this epitome-based compression method is significantly higher than the one obtained with H.264 Intra and with state of the art epitome construction method. A bit-rate saving up to 16% comparatively to H.264 Intra is achieved. Safa Chérigui, Christine Guillemot, Dominique Thoreau, Philippe Guillotel, Patrick Pérez |
MMSP | 2 |
| 2011 | Sparse representations for spatial prediction and texture refinement
Aurélie Martin, Jean-Jacques Fuchs, Christine Guillemot, Dominique Thoreau |
J. Vis. Commun. Image Represent. | 3 |
| 2010 | Sparse optimization with directional DCT bases for image compressionabstractThis paper proposes a new compression algorithm based on the directional DCT (DDCT) bases introduced in [1]. We first explain how to extend the DDCT concept to rectangular bases and exploit them to build a set of bases using a bintree segmentation. We then use dynamic programming to select a basis from this set according to a rate-distortion criterion. Comparisons in terms of rate-distortion performance are finally made with the current compression standards JPEG and JPEG2000. Angélique Dremeau, Cédric Herzet, Christine Guillemot, Jean-Jacques Fuchs |
ICASSP | 3 |
| 2010 | Using an exponential power model forwyner ziv video codingabstractThe Laplacian model is the standard distribution for correlation noise estimation at the turbodecoder in Wyner-Ziv coding schemes. In practice, this hypothesis is not always satisfied and, regularly, the estimated model sensibly differs from the error distribution. In this work, we prove that using a model better fitted to the true distribution improves the performances, and we thus propose to use the more general exponential power distribution (EPD) which has never been tested in a distributed video coding context. Gains in rate-distortion over the Laplacian model are illustrated by results on several video sequences, showing that the EPD model outperforms the Laplacian one in off-line (oracle) as well as in on-line (practical implementation) modes. These results also indicate that, in some cases, the online EPD model reduces the bitrate even over the off-line Laplacian model. Thomas Maugey, Jérôme Gauthier, Béatrice Pesquet-Popescu, Christine Guillemot |
ICASSP | 4 |
| 2010 | Approximate nearest neighbors using sparse representationsabstractA new method is introduced that makes use of sparse image representations to search for approximate nearest neighbors (ANN) under the normalized inner-product distance. The approach relies on the construction of a new sparse vector designed to approximate the normalized inner-product between underlying signal vectors. The resulting ANN search algorithm shows significant improvement compared to querying with the original sparse vectors. The system makes use of a proposed transform that succeeds in uniformly distributing the input dataset on the unit sphere while preserving relative angular distances. Joaquin Zepeda, Ewa Kijak, Christine Guillemot |
ICASSP | 3 |
| 2010 | Hidden Markov Model for distributed video codingabstractThis paper addresses the problem of asymmetric distributed coding of correlated binary Hidden Markov Sources, modeled as a Gilbert-Elliott process. The model parameters are estimated with an estimation-decoding Expectation-Maximization algorithm. The rate gain obtained by accounting for the memory of the sources is first assessed theoretically. The method is then shown to improve the PSNR versus rate performance of a Distributed Video Coding system, based on Low-Density Parity-Check codes. Velotiaray Toto-Zarasoa, Aline Roumy, Christine Guillemot |
ICIP | 3 |
| 2010 | Image prediction: Template matching vs. sparse approximationabstractThe paper compares a sparse approximation based spatial texture prediction method with the template matching based prediction. Template matching algorithms have been widely considered for image prediction. These approaches rely on the assumption that the predicted texture contains a similar textural structure with the template in the sense of a simple distance metric between template and candidate. However, in real images, there are more complex textured areas where template matching fails. The basic idea instead is to consider sparse approximation algorithms. The proposed sparse spatial prediction is assessed against the prediction method based on template matching with a static and optimized dynamic templates. The spatial prediction method is then assessed in a coding scheme where the prediction residue is encoded with a coding approach similar to JPEG. Experimental observations show that the proposed method outperforms the conventional template matching based prediction. Mehmet Türkan, Christine Guillemot |
ICIP | 2 |
| 2010 | Spatial intra-prediction based on mixtures of sparse representationsabstractIn this paper, we consider the problem of spatial prediction based on sparse representations. Several algorithms dealing with this problem can be found in the literature. We propose a novel method involving a mixture of sparse representations. We first place this approach into a probabilistic framework and then derive a practical procedure to solve it. Comparisons of the rate-distortion performance show the superiority of the proposed algorithm with regard to other state-of-the-art algorithms. Angélique Dremeau, Mehmet Türkan, Cédric Herzet, Christine Guillemot, Jean-Jacques Fuchs |
MMSP | 4 |
| 2010 | The Iteration-Tuned Dictionary for sparse representationsabstractWe introduce a new dictionary structure for sparse representations better adapted to pursuit algorithms used in practical scenarios. The new structure, which we call an Iteration-Tuned Dictionary (ITD), consists of a set of dictionaries each associated to a single iteration index of a pursuit algorithm. In this work we first adapt pursuit decompositions to the case of ITD structures and then introduce a training algorithm used to construct ITDs. The training algorithm consists of applying a K-means to the (i -1)-th residuals of the training set to thus produce the i-th dictionary of the ITD structure. In the results section we compare our algorithm against the state-of-the-art dictionary training scheme and show that our method produces sparse representations yielding better signal approximations for the same sparsity level. Joaquin Zepeda, Christine Guillemot, Ewa Kijak |
MMSP | 2 |
| 2010 | On a simple derivation of the complementary matching pursuit
Gagan Rath, Christine Guillemot |
Signal Process. | 2 |
| 2010 | Scalable object-based video retrieval in HD video databases
Claire Morand, Jenny Benois-Pineau, Jean-Philippe Domenger, Joaquin Zepeda, Ewa Kijak, Christine Guillemot |
Signal Process. Image Commun. | 6 |
| 2010 | Perceptually-Friendly H.264/AVC Video Coding Based on Foveated Just-Noticeable-Distortion ModelabstractTraditional video compression methods remove spatial and temporal redundancy based on the signal statistical correlation. However, to reach higher compression ratios without perceptually degrading the reconstructed signal, the properties of the human visual system (HVS) need to be better exploited. Research effort has been dedicated to modeling the spatial and temporal just-noticeable-distortion (JND) based on the sensitivity of the HVS to luminance contrast, and accounting for spatial and temporal masking effects. This paper describes a foveation model as well as a foveated JND (FJND) model in which the spatial and temporal JND models are enhanced to account for the relationship between visibility and eccentricity. Since the visual acuity decreases when the distance from the fovea increases, the visibility threshold increases with increased eccentricity. The proposed FJND model is then used for macroblock (MB) quantization adjustment in H.264/advanced video coding (AVC). For each MB, the quantization parameter is optimized based on its FJND information. The Lagrange multiplier in the rate-distortion optimization is adapted so that the MB noticeable distortion is minimized. The performance of the FJND model has been assessed with various comparisons and subjective visual tests. It has been shown that the proposed FJND model can increase the visual quality versus rate performance of the H.264/AVC video coding scheme. Christine Guillemot |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Robust Video Coding Based on Multiple Description Scalar Quantization With Side InformationabstractThis paper addresses the problem of video compression for robust transmission on lossy Internet networks. The approach developed relies on multiple description coding (MDC) principles. Predictive multiple description video coding has already been considered for robust video transmission over lossy channels. However, MDC, when combined with motion-compensated prediction, is known to suffer from predictive mismatch: in presence of losses, the prediction signal available at the decoder may differ from the one used at the encoder. In this paper, we describe a video compression scheme based on MDC and Wyner-Ziv (WZ) coding. The input sequence is structured into groups of pictures which contain one key frame and one WZ frame. Each frame is first transformed with a wavelet transform and the resulting frequency bands are quantized with a multiple description scalar quantizer. Two balanced descriptions of the video input are thus generated. The quantization indexes of the WZ frames are coded with an low-density parity-check accumulate-based Slepian-Wolf (SW) coder. The lateral receivers first decode the received key frame lateral descriptions, and then construct the side information needed for decoding the corresponding WZ frame descriptions by motion-compensated interpolation. In the case where the two WZ data descriptions are received, the central decoder can perform a separate SW decoding of the two sequences of quantization indexes. A joint iterative decoding approach of the two WZ descriptions is also described which improve the central WZ data decoding performance, however, at the expense of increased decoding complexity. The influence of the proposed iterative decoding technique and of the amount of redundancy on the lateral and central rate-distortion performance of the algorithm is studied. Olivier Crave, Béatrice Pesquet-Popescu, Christine Guillemot |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | Multiple description video coding and iterative decoding of LDPCA codes with side informationabstractIn this paper, we propose the use of multiple description coding to increase the robustness of distributed video coding while keeping good rate-distortion performance. The video sequence is structured into key frames and Wyner-Ziv frames. For each type of frame, two descriptions are generated by a multiple description scalar quantizer and sent on a loss-prone channel. When both Wyner-Ziv descriptions are received, they are jointly decoded along with the side information. We investigate the influence of the amount of redundancy and of the iterative decoding of the descriptions on the performance. Olivier Crave, Christine Guillemot, Béatrice Pesquet-Popescu |
ICASSP | 2 |
| 2009 | Sparse approximation with an orthogonal complementary matching pursuit algorithmabstractThis paper presents the orthogonal extension of the recently introduced complementary matching pursuit (CMP) algorithm for sparse approximation. The CMP algorithm is analogous to the matching pursuit (MP) but done in the row-space of the dictionary matrix. It suffers from a similar sub-optimality as the MP. The orthogonal complementary matching pursuit algorithm (OCMP) presented here tries to remove this sub-optimality by updating the coefficients of all selected atoms at each iteration. Its development from the CMP follows the same procedure as of the orthogonal matching pursuit (OMP). In contrast with OMP, the residual errors resulting from the OCMP may not be orthogonal to all the atoms selected up to the respective iteration. Though the residual energy may increase over the OMP during the first iterations, it is shown that, compared with OMP, the convergence speed is increased in the subsequent iterations and the sparsity of the solution vector is improved. Gagan Rath, Christine Guillemot |
ICASSP | 2 |
| 2009 | Robust and fast non asymmetric distributed source coding using turbo codes on the syndrome trellisabstractWe consider the distributed compression of two (binary memoryless) correlated sources and propose a unique codec that can reach any point in the Slepian-Wolf region. In a previous method based on channel codes, the decoder multiply the compressed data by an inverse submatrix of the code. This multiplication presents two drawbacks. First, if turbo codes are used, the submatrix has no periodic structure s.t. the whole inverse has to be stored and no fast implementation exists for the multiplication. Second, this multiplication may lead to error propagation. In this paper, we propose a method that is both robust and fast. Velotiaray Toto-Zarasoa, Aline Roumy, Christine Guillemot, Cédric Herzet |
ICASSP | 3 |
| 2009 | Perceptually-friendly H.264/AVC video codingabstractThis paper presents a perceptual H.264/AVC video coding method based on a foveated just-noticeable-distortion (JND) model. Since the perceptual acuity decreases with the increased eccentricity, a foveation model is adopted to further explore the perceptual redundancy in addition to the spatial and temporal just-noticeable-distortion models. Bit allocation and rate-distortion optimization algorithms based on the foveated model are addressed. The performance of the foveated JND model is assessed with subjective visual tests. Applying the proposed model in H.264/AVC video coding can achieve better visual quality. Christine Guillemot |
ICIP | 2 |
| 2009 | Sparse approximation with adaptive dictionary for image predictionabstractThe paper presents a dictionary construction method for spatial texture prediction based on sparse approximations. Sparse approximations have been recently considered for image prediction using static dictionaries such as a DCT or DFT dictionary. These approaches rely on the assumption that the texture is periodic, hence the use of a static dictionary formed by pre-defined waveforms. However, in real images, there are more complex and non-periodic textures. The main idea underlying the proposed spatial prediction technique is instead to consider a locally adaptive dictionary, A, formed by atoms derived from texture patches present in a causal neighborhood of the block to be predicted. The sparse spatial prediction method is assessed against the sparse prediction method based on a static DCT dictionary. The spatial prediction method is then assessed in a complete image coding scheme where the prediction residue is encoded using a coding approach similar to JPEG. Mehmet Türkan, Christine Guillemot |
ICIP | 2 |
| 2009 | SIFT-based local image description using sparse representationsabstractThis paper addresses the problem of efficient SIFT-based image description and searches in large databases within the framework of local querying. A descriptor called the bag-of-features has been introduced previously which first vector quantizes SIFT descriptors and then aggregates the set of resulting codeword indices (so-called visual words) into a histogram of occurrence of the different visual words in the image. The aim is to make the image search complexity tractable by transforming the set of local image descriptor vectors into a single sparse vector as sparsity particularly permits efficient inner product calculations. However, aggregating local descriptors into a single histogram decreases the discerning power of the system when performing local queries. In this paper, we propose a new approach that aims to enjoy the complexity benefits of sparsity while at the same time retaining the local quality of the input descriptor vectors. This is accomplished by searching for a sparse approximation of the input SIFT descriptors. The sparse approximation yields a sparse vector per local SIFT descriptor, and helps preserving local description properties by using each sparse-transformed descriptor independently in a voting system to retrieve indexed images. Our system is shown experimentally to perform better than histogram based systems under query locality, albeit at an increased complexity. Joaquin Zepeda, Ewa Kijak, Christine Guillemot |
MMSP | 3 |
| 2009 | Perception-oriented video coding based on foveated JND modelabstractThe ultimate goal of video compression is to provide best visual quality of the video source by using minimum bit budget. To achieve this target, spatial, temporal and perceptual redundancies should be exploited and removed such that the video source could be compressed efficiently. In this paper, we present how to improve the video compression performance by using a foveated just-noticeable-distortion model. Since the perceptual acuity decreases with the increased eccentricity, a foveation model is developed to exploit the global perceptual redundancy in addition to the spatial and temporal just-noticeable-distortion models. The foveated just-noticeable-distortion model is incorporated into the H.264/AVC video coding system. As demonstrated in the subjective tests, applying the proposed model in video coding can improve visual quality. Christine Guillemot |
PCS | 2 |
| 2009 | Distributed coding using punctured quasi-arithmetic codes for memory and memoryless sourcesabstractThis paper considers the use of punctured quasi-arithmetic (QA) codes for the Slepian-Wolf problem. These entropy codes are defined by finite state machines for memory-less and first-order memory sources. Puncturing an entropy coded bit-stream leads to an ambiguity at the decoder side. The decoder makes use of a correlated version of the original in order to remove this ambiguity. A complete DSC scheme based on QA encoding with side information at the decoder is presented. The proposed scheme is adapted to memoryless and first-order memory sources. Simulation results reveal that the proposed scheme is efficient in terms of decoding performance for short sequences compared to well-known DSC using channel codes. Simon Malinowski, Xavier Artigas, Christine Guillemot |
PCS | 3 |
| 2009 | Computation of posterior marginals on aggregated state models for soft source decodingabstractOptimum soft decoding of sources compressed with variable length codes and quasi-arithmetic codes, transmitted over noisy channels, can be performed on a bit/symbol trellis. However, the number of states of the trellis is a quadratic function of the sequence length leading to a decoding complexity which is not tractable for practical applications. The decoding complexity can be significantly reduced by using an aggregated state model, while still achieving close to optimum performance in terms of bit error rate and frame error rate. However, symbol a posteriori probabilities can not be directly derived on these models and the symbol error rate (SER) may not be minimized. This paper describes a two-step decoding algorithm that achieves close to optimal decoding performance in terms of SER on aggregated state models. A performance and complexity analysis of the proposed algorithm is given. Simon Malinowski, Hervé Jégou, Christine Guillemot |
IEEE Trans. Commun. | 3 |
| 2008 | Rate-adaptive codes for the entire Slepian-Wolf region and arbitrarily correlated sourcesabstractIn this paper, we focus on the design of distributed source codes that can achieve any point in the Slepian-Wolf (SW) region and at the same time adapt to any correlation between the sources. A practical solution based on punctured accumulated LDPC codes extended to the non asymmetric case is described. The approach allows flexible rate allocation to the two sources with a gap of 0.0677 bits with respect to the minimum achievable rate. Velotiaray Toto-Zarasoa, Aline Roumy, Christine Guillemot |
ICASSP | 3 |
| 2008 | Atomic decomposition dedicated to AVC and spatial SVC predictionabstractThis work, we propose the use of sparse signal representation techniques to solve the problem of closed-loop spatial image prediction. The reconstruction of signal in the block to predict is based on basis functions selected with the matching pursuit (MP) iterative algorithm, to best match a causal neighborhood. We evaluate this new method in terms of PSNR and bitrate in a H.264/AVC encoder. Experimental results indicate an improvement of rate-distortion performance. In this paper, we also present results concerning the use of this technique for intra-inter layer prediction refinement, in a scalable video coding (SVC) like scheme. Aurélie Martin, Jean-Jacques Fuchs, Christine Guillemot, Dominique Thoreau |
ICIP | 3 |
| 2008 | Spatial prediction in distributed video coding using Wyner-Ziv DPCMabstractThis paper deals with spatial prediction in Wyner-Ziv video coding. Coding structures that extend DPCM to the distributed coding scenario are presented. Using a high resolution approximation, the optimal prediction filters for these structures are derived. The resulting predictive coder is employed to compress the DC component of a transform based distributed video coding system. Preliminary results show rate-distortion gains compared to the case without spatial prediction. Jayanth Nayak, Gagan Rath, Denis Kubasov, Christine Guillemot |
MMSP | 4 |
| 2008 | Sparse approximations for joint source-channel codingabstractThis paper considers the application of sparse approximations in a joint source-channel (JSC) coding framework. The considered JSC coded system employs a real number BCH code on the input signal before the signal is quantized and further processed. Under an impulse channel noise model, the decoding of error is posed as a sparse approximation problem. The orthogonal matching pursuit (OMP) and basis pursuit (BP) algorithms are compared with the syndrome decoding algorithm in terms of mean square reconstruction error. It is seen that, with a Gauss-Markov source and Bernoulli-Gaussian channel noise, the BP outperforms the syndrome decoding and the OMP at higher noise levels. In the case of image transmission with channel bit errors, the BP outperforms the other two decoding algorithms consistently. Gagan Rath, Christine Guillemot, Jean-Jacques Fuchs |
MMSP | 2 |
| 2008 | Special issue on distributed video coding
Christine Guillemot, Fernando Pereira 0001 |
Signal Process. Image Commun. | 1 |
| 2008 | Distributed Video Coding: Selecting the most promising application scenarios
Fernando Pereira 0001, Christine Guillemot, Touradj Ebrahimi, Riccardo Leonardi, Sven Klomp |
Signal Process. Image Commun. | 3 |
| 2008 | Compression of Laplacian Pyramids Through Orthogonal Transforms and Improved PredictionabstractScalable representation of visual signals, such as image and video signals, has become a subject of active research since early 1980s. Scalability allows the adaptation of the bit rate and/or the resolution of the transmitted data to the network bandwidth and/or the rendering capability of the receiving device. For many years, spatial scalability has been achieved through wavelets, but recently the Laplacian pyramid (LP) has become an alternative choice because of reduced aliasing in the lower resolutions. In this paper, we focus on the coding efficiency of the LP with a view to transmitting it over a communication channel. In particular, we aim to improve the compression efficiency of the LP detail layers through improved interlayer prediction and orthogonal spatial transforms. First, we consider an LP in the open-loop configuration and propose to improve its rate-distortion performance by compressing it to a critically sampled representation. We derive four different orthogonal spatial transforms from the upsampling and downsampling filters that can achieve this representation, and apply them on the detail layers. The application of these transforms to the detail layers renders a fixed number of transform coefficients either zero or redundant, thus making their transmission unnecessary. Then we consider the compression of an LP in the closed-loop configuration through similar spatial transforms. Because of the introduction of quantization in the prediction loop, these spatial transforms applied on the detail layers do not produce the same number of zero or redundant transform coefficients as in the open-loop case. Nevertheless, the insight obtained from the open-loop coding leads us to enhance the interlayer prediction, and the subsequent application of the spatial transforms to the new detail layers aims to achieve better energy compaction. Gagan Rath, Christine Guillemot |
IEEE Trans. Image Process. | 3 |
| 2007 | Optimal Distributed Linear Transceivers for Sending Independently Corrupted Copies of a Colored Source Over the Gaussian MACabstractWe consider distributed linear transceivers for sending a second-order wide-sense stationary process observed by two noisy sensors over a Gaussian multiple-access channel (MAC). We derive the minimum mean-square error (MSE) distributed linear transceiver. The optimal linear transmitter exploits bandwidth expansion by repeating transmission and the transmitters at the two sensors are the same except for a constant factor. When the source is white, encoded transmission is the best linear code for any SNR. But for a colored source, whitening transmit filter is sub-optimal. In high SNR regime, the magnitude response of the optimal transmission filter is inversely proportional to fourth-root of the power spectrum of the process (while that for the whitening filler is inversely proportional to the square-root of the spectrum). In the special case of n single sensor with Gaussian source, we also quantify the performance loss of linear source-channel codes with respect to the Shannon limit. Onkar Dabeer, Aline Roumy, Christine Guillemot |
ICASSP (3) | 3 |
| 2007 | Improved Interlayer Prediction for Scalable Video CodingabstractIn this paper, we present a novel technique for efficient compression of enhancement layers in scalable video coding (SVC). First we propose an improved interlayer prediction scheme which exploits the inherent redundancy of the underlying Laplacian pyramid with nonbiorthogonal filters. Secondly we introduce an orthogonal transform in parallel with the current 4×4 transform to improve the coding efficiency of the enhancement layer further. The improved prediction and the transform are implemented in the SVC reference software JSVM 4.0 as additional prediction modes. Experimental results demonstrate coding gains up to 1 dB for I pictures, and up to 0.7 dB for both I and P pictures, over a current implementation. Gagan Rath, Christine Guillemot |
ICASSP (1) | 3 |
| 2007 | Overlapped Quasi-Arithmetic Codes for Distributed Video CodingabstractThis paper describes Slepian-Wolf codes based on overlapped quasi-arithmetic codes, where overlapping allows lossy compression of the source below its entropy. In the context of separate decoding, these codes are not uniquely decodable: the overlap introduces ambiguity in the decoding process leading to decoding errors. The presence of correlated side information at the decoder is used to remove this ambiguity and achieve a vanishing error probability. The state models and the automata of the overlapped quasi-arithmetic codes are described. The soft decoding algorithm with side information is then presented. The performance of these codes has been assessed first on theoretical sources and integrated in a distributed video coding platform. Xavier Artigas, Simon Malinowski, Christine Guillemot |
ICIP (2) | 3 |
| 2007 | A Hybrid Encoder/Decoder Rate Control for Wyner-Ziv Video Coding with a Feedback ChannelabstractThis paper describes a hybrid coder/decoder rate control for a Wyner-Ziv video coding scheme with a feedback channel. The approach is first based on a method to estimate at the encoder the minimum rate required for the Slepian-Wolf (SW) encoded data. This estimation makes use only of the Lapacian correlation model. The decoder then estimates the bit error rate (BER) for each decoded bit plane from likelihood ratios computed at the output of the SW decoder. The robustness of the BER estimation is further improved by the use of an error detection mechanism based on a cyclic redundancy checksum. Experimental results show that rate-distortion performances are comparable to those which would be obtained by computing the Hamming distance between the original and the decoded data. In addition the hybrid coder/decoder rate control reduces the decoding complexity as well as the usage of the return channel. Denis Kubasov, Khaled Lajnef, Christine Guillemot |
MMSP | 3 |
| 2007 | Optimal Reconstruction in Wyner-Ziv Video Coding with Multiple Side InformationabstractThis paper addresses the problem of optimal minimum mean-squared error reconstruction of quantised samples in Wyner-Ziv video coding systems. Closed-form expressions of the optimal reconstructed values are derived for a Laplacian correlation model. The method is used for both single and multiple side information scenarios (the latter is also referred to as multi-hypothesis Wyner-Ziv decoding). The efficiency of the proposed optimal reconstruction method is confirmed by rate-distortion performance results, showing significant decrease of the distortion of the decoded sequence, compared to simple reconstruction methods that have been employed so far. Denis Kubasov, Jayanth Nayak, Christine Guillemot |
MMSP | 3 |
| 2007 | Improved Intra Prediction for H.264/AVC Scalable ExtensionabstractThis paper presents an improved intra prediction scheme for H.264/AVC scalable extension. First, a more flexible intra prediction is realized by introducing a sub-macroblock (sub-MB) level inter-layer intra prediction. Second, the upsampled reconstructed base layer helps to reduce the number of candidate modes for the spatial intra prediction. Finally, the neighboring pixels of the reconstructed enhancement layer are used to improve the quality of the reference from the base layer by a simple cross boundary Alter. Our scheme improves the PSNR up to 0.53 dB when compared with the JSVM standard implementation for all intra coding. Gagan Rath, Christine Guillemot |
MMSP | 3 |
| 2007 | Joint Source-Channel Turbo Techniques for Discrete-Valued Sources: From Theory to PracticeabstractThe principles which have been prevailing so far for designing communication systems rely on Shannon's source and channel coding separation theorem. This theorem states that source and channel optimum performance bounds can be approached as close as desired by designing independently the source and channel coding strategies. However, this theorem holds only under asymptotic conditions, where both codes are allowed infinite length and complexity. If the design of the system is constrained in terms of delay and complexity, if the sources are not stationary, or if the channels are nonergodic, separate design and optimization of the source and channel coders can be largely suboptimal. For practical systems, joint source–channel (de)coding may reduce the end-to-end distortion. It is one of the aspects covered by the term cross-layer design, meaning a rethinking of the layer separation principle. This article focuses on recent developments of joint source–channel turbo coding and decoding techniques, which are described in the framework of normal factor graphs. The scope is restricted to lossless compression and discrete-valued sources. The presented techniques can be applied to the quantized values of a lossy source codec but the quantizer itself and its impact are not considered. Xavier Jaspar, Christine Guillemot, Luc Vandendorpe |
Proc. IEEE | 2 |
| 2007 | Entropy Coding With Variable-Length Rewriting SystemsabstractThis paper describes a family of codes for entropy coding of memoryless sources. These codes are defined by sets of production rules of the form $a\bar{l} \rightarrow\bar{b}$ , where $a$ is a source symbol, and $\bar{l},\bar{b}$ are sequences of bits. The coding process can be modeled as a finite-state machine (FSM). A method to construct codes which preserve the lexicographic order in the binary-coded representation is described. For a given constraint on the number of states for the coding process, this method allows the construction of codes with a better compression efficiency than the HuTucker codes. A second method is proposed to construct codes such that the marginal bit probability of the compressed bitstream converges to 0.5 as the sequence length increases. This property is achieved even if the probability distribution function is not known by the encoder. Hervé Jégou, Christine Guillemot |
IEEE Trans. Commun. | 2 |
| 2007 | 3-D Model-Based Frame Interpolation for Distributed Video Coding of Static ScenesabstractThis paper addresses the problem of side information extraction for distributed coding of videos captured by a camera moving in a 3-D static environment. Examples of targeted applications are augmented reality, remote-controlled robots operating in hazardous environments, or remote exploration by drones. It explores the benefits of the structure-from-motion paradigm for distributed coding of this type of video content. Two interpolation methods constrained by the scene geometry, based either on block matching along epipolar lines or on 3-D mesh fitting, are first developed. These techniques are based on a robust algorithm for sub-pel matching of feature points, which leads to semi-dense correspondences between key frames. However, their rate-distortion (RD) performances are limited by misalignments between the side information and the actual Wyner-Ziv (WZ) frames due to the assumption of linear motion between key frames. To cope with this problem, two feature point tracking techniques are introduced, which recover the camera parameters of the WZ frames. A first technique, in which the frames remain encoded separately, performs tracking at the decoder and leads to significant RD performance gains. A second technique further improves the RD performances by allowing a limited tracking at the encoder. As an additional benefit, statistics on tracks allow the encoder to adapt the key frame frequency to the video motion content. Matthieu Maitre, Christine Guillemot, Luce Morin |
IEEE Trans. Image Process. | 2 |
| 2007 | Synchronization Recovery and State Model Reduction for Soft Decoding of Variable Length CodesabstractVariable length codes (VLCs) exhibit loss of synchronization problems when transmitted over noisy channels. Trellis decoding techniques based on Maximum A Posteriori (MAP) estimators are often used to minimize the error rate on the estimated sequence. If the number of symbols and/or bits transmitted is known by the decoder, termination constraints can be incorporated in the decoding process. All the paths in the trellis which do not lead to a valid sequence length are suppressed. This correspondence presents an analytic method to assess the expected error resilience of a VLC when trellis decoding with a sequence length constraint is used. The approach is based on the computation, for a given code, of the amount of information brought by the constraint. It is then shown that this quantity is not significantly altered by appropriate trellis states aggregation. This proves that the performance obtained by running a length-constrained Viterbi decoder on aggregated state models approaches the one obtained with the bit/symbol trellis, with a significantly reduced complexity. It is then shown that the complexity can be further decreased by projecting the state model on two state models of reduced size. Simon Malinowski, Hervé Jégou, Christine Guillemot |
IEEE Trans. Inf. Theory | 3 |
| 2006 | Mesh-Based Motion-Compensated Interpolation for Side Information Extraction in Distributed Video CodingabstractThis paper addresses the problem of side information generation in distributed video compression (DVC) schemes. Intermediate frames constructed by motion-compensated interpolation of key frames are used as side information to decode Wyner-Ziv frames. The limitations of block-based translational motion models call for new motion models. This article studies the benefits of mesh-based motion-compensated (MC) interpolation for side information extraction in DVC. A hybrid block-based and mesh-based solution addressing the problem of motion discontinuities and occlusions is also described. The increased correlation between the side information and the Wyner-Ziv encoded frames leads to performance gains which can reach up to 1 dB with respect to block-based solutions. Denis Kubasov, Christine Guillemot |
ICIP | 2 |
| 2006 | 3D Scene Modeling for Distributed Video CodingabstractThe compression efficiency of distributed video-coding (DVC) suffers from the necessity of transmitting a large number of key-frames which are intra-coded. This paper describes a new 3D model-based DVC approach which reduces the key- frame frequency. The decoder first recovers a 3D model from the key-frames. It then predicts the intermediate frames by projecting it onto 2D image planes and applying image-based rendering techniques. This paper also introduces a new quasi-DVC method relying on a limited point tracking at the encoder. It greatly improves the prediction PSNR, while only slightly increasing the encoder complexity. It also allows the encoder to adaptively select the key-frames based on the video motion-content. Matthieu Maitre, Christine Guillemot, Luce Morin |
ICIP | 2 |
| 2006 | An SVD Based Transform for Critical Representation of Laplacian PyramidsabstractRecently the Laplacian pyramid (LP) has come to prominence as a tool for obtaining spatial scalability in the context of scalable image and video compression. In this paper, we present a critical representation scheme for the LP. The representation is obtained by transforming the detail signals to lower dimensional representations. Unlike the wavelet transform, the proposed transform depends on the decimation and the interpolation filters used for the classical LP. Simulation results suggest that higher compression ratios can be achieved with the critical representation than with the standard LP with usual or frame based reconstructions. Gagan Rath, Christine Guillemot |
ICIP | 2 |
| 2006 | Error recovery properties of quasi-arithmetic codes and soft decoding with length constraintabstractIn this paper, we propose a method to analyse the error recovery properties of quasi-arithmetic codes. This method is adapted from the one proposed in J. Maxted and J. Robinson (1985) for variable length codes. The expected number of symbols affected by a single bit error and the probability mass function of the gain/loss (P.F. Swaszek and P. DiCicco, 1995) of symbols following a single bit error can be computed with this method. A method to estimate this probability mass function when a bitstream is sent over a binary symmetrical channel is then proposed. The aggregated state model for soft decoding of variable length codes proposed in H. Jegou et al. (2005) is then extended to quasi-arithmetic codes, as the synchronisation recovery properties of both kind of codes are similar. The soft decoding results of this scheme reveal high performance with a reasonable computing cost Simon Malinowski, Hervé Jégou, Christine Guillemot |
ISIT | 3 |
| 2006 | Compressing the Laplacian PyramidabstractThe Laplacian pyramid (LP) is one of the earliest examples of multiscale representation of visual data. It is well known that an LP is overcomplete or redundant by construction, and has lower compression efficiency compared to critical representations such as wavelets and subband coding. In this paper, we propose to improve the rate-distortion (R-D) performance of the LP through critical representation. We consider an LP with biorthogonal decimation and interpolation filters, and show that the detail signals lie in lower-dimensional subspaces. This allows them to be represented using fewer coefficients than the original spatial representations. We derive orthogonal bases for these subspaces and represent the detail signals in terms of their projections onto these bases. Simulation results suggest that higher compression ratios can be achieved with the critical representation than with the standard LP with usual or dual frame based reconstructions Gagan Rath, Christine Guillemot |
MMSP | 2 |
| 2006 | Distributed coding of three binary and Gaussian correlated sources using punctured Turbo Codes
Khaled Lajnef, Christine Guillemot, Pierre Siohan |
Signal Process. | 2 |
| 2006 | Oriented Wavelet Transform for Image Compression and DenoisingabstractIn this paper, we introduce a new transform for image processing, based on wavelets and the lifting paradigm. The lifting steps of a unidimensional wavelet are applied along a local orientation defined on a quincunx sampling grid. To maximize energy compaction, the orientation minimizing the prediction error is chosen adaptively. A fine-grained multiscale analysis is provided by iterating the decomposition on the low-frequency band. In the context of image compression, the multiresolution orientation map is coded using a quad tree. The rate allocation between the orientation map and wavelet coefficients is jointly optimized in a rate-distortion sense. For image denoising, a Markov model is used to extract the orientations from the noisy image. As long as the map is sufficiently homogeneous, interesting properties of the original wavelet are preserved such as regularity and orthogonality. Perfect reconstruction is ensured by the reversibility of the lifting scheme. The mutual information between the wavelet coefficients is studied and compared to the one observed with a separable wavelet transform. The rate-distortion performance of this new transform is evaluated for image coding using state-of-the-art subband coders. Its performance in a denoising application is also assessed against the performance obtained with other transforms or denoising methods. Vivien Chappelier, Christine Guillemot |
IEEE Trans. Image Process. | 2 |
| 2005 | Oriented wavelet transform on a quincunx pyramid for image compressionabstractThis paper presents a new image coding technique based on oriented ID multiscale decompositions on a quincunx sampling grid. A transform obtained by adapting the lifting steps of a ID wavelet transform along a local orientation is described. A quad-tree structure is used to efficiently code the orientation map, and optimized in the rate-distortion sense. A gain of up to 1.3 dB over the separable wavelet transform is observed for a set of images, with the rate estimated from the stationary entropy of the subbands. Vivien Chappelier, Christine Guillemot |
ICIP (1) | 2 |
| 2005 | Entropy coding with variable length re-writing systemsabstractThis paper describes a new set of block source codes well suited for data compression. These codes are defined by sets of productions rules of the form allowbar rarr blowbar where a isin A represents a value from the source alphabet A and llowbar, blowbar are small sequences of bits. These codes naturally encompass other variable length codes (VLCs) such as Huffman codes. It is shown that these codes may have a similar or even a shorter mean description length than Huffman codes for the same encoding and decoding complexity. A first code design method allowing to preserve the lexicographic order in the bit domain is described. The corresponding codes have the same mean description length (mdl) as Huffman codes from which they are constructed. Therefore, they outperform from a compression point of view the Hu-Tucker codes designed to offer the lexicographic property in the bit domain. A second construction method allows to obtain codes such that the marginal bit probability converges to 0.5 as the sequence length increases and this is achieved even if the probability distribution function is not known by the encoder Hervé Jégou, Christine Guillemot |
ISIT | 2 |
| 2005 | Robust multiplexed codes for compression of heterogeneous dataabstractCompression systems of real signals (images, video, audio) generate sources of information with different levels of priority which are then encoded with variable-length codes (VLCs). This paper addresses the issue of robust transmission of such VLC encoded heterogeneous sources over error-prone channels. VLCs are very sensitive to channel noise: when some bits are altered, synchronization losses can occur at the receiver. This paper describes a new family of codes, called multiplexed codes, that confine the de-synchronization phenomenon to low-priority data while reaching asymptotically the entropy bound for both (low- and high-priority) sources. The idea consists of creating fixed-length codes for high-priority information and of using the inherent redundancy to describe low-priority data, hence the name multiplexed codes. Theoretical and simulation results reveal a very high error resilience at almost no cost in compression efficiency. Hervé Jégou, Christine Guillemot |
IEEE Trans. Inf. Theory | 2 |
| 2004 | Joint Source-Channel Decoding of Quasi-Arithmetic CodesabstractThis paper addresses the issue of robust and joint source-channel decoding of quasiarithmetic codes. Quasiarithmetic coding is a reduced precision and complexity implementation of arithmetic coding. This paper provides first a state model of a quasiarithmetic decoder for binary and M-ary sources. The design of an error-resilient soft decoding algorithm follows quite naturally. The compression efficiency of quasiarithmetic codes allows to add extra redundancy in the form of markers designed specifically to prevent de-synchronization. The algorithm is directly amenable for iterative source-channel decoding in the spirit of serial turbo codes. The coding and decoding algorithms have been tested for a wide range of channel signal-to-noise ratios. Experimental results reveal improved SER and SNR performances against Huffman and optimal arithmetic codes. Thomas Guionnet, Christine Guillemot |
Data Compression Conference | 2 |
| 2004 | Joint source and channel coding/decoding by concatenating an oversampled filter bank code and a redundant entropy code [image coding application]abstractThis paper examines a joint source and channel coding (JSCC) technique based on oversampled filter banks (OFB). We consider an image compression system with an OFB based signal decomposition followed by scalar quantization and a redundant entropy coding scheme. This encoding chain can be viewed as a concatenation of an OFB code and an error correcting entropy code. The error correcting ability of an OFB comes from the explicit redundancy introduced into the transmitted signal by oversampling while the error correcting ability of the entropy code comes from the implicit redundancy due to the Markov property of the OFB outputs. The decoder consists of the soft-input soft-output entropy decoder and a syndrome decoder for the OFB code. The soft information at the output of the entropy decoder is utilized in the error localization procedure of the OFB syndrome decoder. The performance of the algorithm is examined for an image coding system with a wavelet filter bank. Slavica Marinkovic, Christine Guillemot |
GLOBECOM | 2 |
| 2004 | Erasure resilience of oversampled filter bank codes based on cosine modulated filter banksabstractThis paper examines erasure resilience of oversampled filter bank (OFB) codes focusing on two families of codes based on cosine modulated filter banks (CMFB). We discuss frame properties of the CMFB based OFB codes and prove perfect reconstruction (PR) for erasure patterns for which PR depends only on the general structure of the code and not on the prototype filter coefficients. For some of these erasure patterns the expression for the mean square reconstruction error is also independent of filter coefficients and can be expressed in terms of the number of erasures, and of parameters such as number of channels and oversampling ratio. We further take an image coding system as an example and verify the performance of a particular OCMFB code for all possible erasure patterns by simulation. Slavica Marinkovic, Christine Guillemot |
ICC | 2 |
| 2004 | Accuracy-scalable motion coding for efficient scalable video compressionabstractFor a scalable video coder to remain efficient over a wide range of bit-rates, covering e.g. both mobile video streaming and TV broadcasting, some form of scalability must exist in the motion information. In this paper we propose a new (1+2D) wavelet-based spatio-SNR-temporal-scalable video codec, coupled with an accuracy-scalable motion codec. It allows to decode a reduced amount of motion information at subresolutions, taking advantage that motion compensation requires less and less accuracy at lower spatial resolutions. This new motion codec proves its efficiency in our full-scalable framework, by improving significantly video quality at subresolutions without inducing any noticeable penalty at high bit-rates. Guillaume Boisson, Edouard François, Christine Guillemot |
ICIP | 3 |
| 2004 | Image coding with iterated contourlet and wavelet transformsabstractThis paper presents a new coding technique based on a mixed contourlet and wavelet transform. The redundancy of the transform is controlled by using the contourlet at fine scales and by switching to a separable wavelet transform at coarse scales. The transform is then optimized through an iterative projection process in the transform domain in order to minimize the quantization error in the image domain. A gain of respectively up to 0.5 dB and to 1 dB over respectively contourlet and wavelet based coding has been observed for images with directional features. Vivien Chappelier, Christine Guillemot, Slavica Marinkovic |
ICIP | 2 |
| 2004 | A para-pseudo inverse based method for reconstruction of filter bank frame-expanded signals from erasuresabstractPacket losses due to congestion or buffer overflows is a common problem in packet switched networks. Current network protocols manage this problem by retransmitting the lost packets. However, the delay due to the retransmission of the lost packets may be unacceptable for many real-time applications. Recent focus to resolve this problem is to recover the lost data from the received packets using some error control coding scheme. In this context, signal representation using frames has gained attention and has been studied in J. Kovacevic et al. (2002), P.L. Dragotti et al. (2001), G. Rath and C. Guillemot (2004), and R. Motwani and C. Guillemot (2004). Oversampled transforms and oversampled filter banks have been considered as joint-source channel codes and methods for reconstructing from erasures is studied in these articles. However, for oversampled filter banks, the reconstruction methods based on operating the pseudo-inverse based on the entire signal length are computationally complex and those based on reconstructing the erasures are not optimal as far as reconstruction mean square error is concerned. In this paper, we propose a method for reconstruction from erasures using a synthesis filter bank which functions as a pseudo-inverse. Hence, the scheme minimizes the reconstruction mean square error. Further, the method is computationally efficient, because it does not operate the pseudo-inverse corresponding to the entire signal vector. The synthesis filter bank, which obviously depends on the erasure pattern implements the pseudo-inverse at a practical computational cost. Some typical bursty erasure patterns which permit existence of a FIR synthesis filter banks are studied. The theoretical results are validated for bursty erasure patterns by simulations using image data. Ravi Motwani, Christine Guillemot |
ICIP | 2 |
| 2004 | First-order multiplexed source codes for error-resilient entropy codingabstractThis paper describes a new class of variable length codes (VLCs) that allow to exploit first-order source statistics while still being resilient to transmission errors. This paper extends the work of [H. Jegou et al., (2003)] to take into account the source conditional probabilities. Theoretical performances in terms of compression efficiency and error resilience are analyzed. Hervé Jégou, Christine Guillemot |
ISIT | 2 |
| 2004 | Filter bank frame-expansion with erasures-a para-pseudo inverse based reconstructionabstractIn this paper, we propose a decoder which reconstructs signals which are joint-source channel encoded using oversampled filter banks. The decoder is based on operating the para-pseudoinverse corresponding to the cascaded transform function consisting of the source encoder and the erasure channel on the received signal. Hence, the proposed method minimizes the reconstruction mean square error. The decoder consists of a filter bank which can comprise of FIR or IIR filters depending on the erasure channel. Hence the computational complexity of the decoder is also lower compared to earlier proposed techniques. The theoretical results are validated by simulations using image data. Ravi Motwani, Christine Guillemot |
ISIT | 2 |
| 2004 | Distributed coding of three sources using punctured turbo codesabstractThis paper describes a distributed coding system based on punctured turbo codes to compress three correlated memoryless binary and Gaussian sources. The performance bounds given by the Slepian-Wolf and Wyner-Ziv theorems are provided for three sources. The simulation results are analyzed against those provided by the same asymmetric system based on punctured turbo codes used for distributed coding of two sources. Khaled Lajnef, Christine Guillemot, Pierre Siohan |
MMSP | 2 |
| 2004 | Joint source-channel coding as an element of a QoS framework for '4G' wireless multimedia
Christine Guillemot, Paul Christ 0001 |
Comput. Commun. | 1 |
| 2004 | Subspace algorithms for error localization with quantized DFT codesabstractRecently, a class of real-number Bose-Chaudhuri-Hocquengem codes known as discrete Fourier transform (DFT) codes have been considered as joint source and channel codes for providing robustness to erasures and errors over wireless networks. We propose three subspace algorithms for error localization with quantized DFT codes. The algorithms are similar to the MUSIC, the minimum-norm, and the ESPRIT algorithms used in array signal processing for direction-of-arrival estimation. They provide different but related formulations of the error localizations by first partitioning a vector space into the channel error subspace and its orthogonal complement, the noise subspace. The locations of the errors are determined from either the error subspace eigenvectors or the noise subspace eigenvectors. We also present a brief performance analysis of the localization error in terms of the perturbation of the error subspace due to quantization. Simulation results show that their localization performances are similar, and they perform better than the coding-theoretic approach over a broad range of channel-error-to-quantization-noise ratios. Gagan Rath, Christine Guillemot |
IEEE Trans. Commun. | 2 |
| 2004 | Real-time constrained TCP-compatible rate control for video over the InternetabstractThis paper describes a rate control algorithm that captures not only the behavior of TCP's congestion control avoidance mechanism but also the delay constraints of real-time streams. Building upon the TFRC protocol , a new protocol has been designed for estimating the bandwidth prediction model parameters. Making use of RTP and RTCP, this protocol allows to better take into account the multimedia flows characteristics (variable packet size, delay ...). Given the current channel state estimated by the above protocol, encoder and decoder buffers states as well as delay constraints of the real-time video source are translated into encoder rate constraints. This global rate control model, coupled with an H.263+ loss resilient video compression algorithm, has been extensively experimented with on various Internet links. The experiments clearly demonstrate the benefits of 1/ the new protocol used for estimating the bandwidth prediction model parameters, adapted to multimedia flows characteristics, and of 2/ the global rate control model encompassing source buffers and end-to-end delay characteristics. The overall system leads to reduce significantly the source timeouts, hence to minimize the expected distortion, for a comparable usage of the TCP-compatible predicted bandwidth. Jérôme Viéron, Christine Guillemot |
IEEE Trans. Multim. | 2 |
| 2003 | ESPRIT-like error localization algorithm for a class of real number codesabstractJoint source and channel coding of multimedia data for transmission over wireless networks is a challenging task. Recently, a class of BCH-like real number codes has been considered for providing robustness to data loss or data error over such networks. Unlike finite-field codes, the real codevectors undergo a quantization process before being transmitted, and therefore, the error localization efficiency becomes algorithm-dependent. The paper presents an ESPRIT-like subspace algorithm for localizing errors in such codes and compares its performance with other known localization algorithms. In particular, we consider lowpass DFT, DCT and DST codes, and compare the localization efficiency of the ESPRIT-like algorithm with the standard BCH decoding approach and the MUSIC-like subspace approach. Simulation results with a Gauss-Markov source are presented. Gagan Rath, Christine Guillemot |
GLOBECOM | 2 |
| 2003 | Error-resilient binary multiplexed source codesabstractThe paper addresses the issue of robust transmission of VLC encoded sources over error-prone channels. We have recently introduced a new family of codes, called multiplexed codes. They exploit the fact that real signal compression systems generate sources of information with different levels of priority. Multiplexed codes allow the desynchronization phenomenon to be confined to low priority data while allowing the entropy bound to be reached asymptotically for both (low and high priority) sources. A multiplexing procedure based on an iterative Euclidian decomposition has been proposed. This paper introduces a variant of multiplexed codes, called binary multiplexed codes, together with a very simple multiplexing algorithm that exploits the structure of variable length codetrees. It is shown analytically and experimentally that this family of codes is more error resilient than fixed length codes while reaching the compression efficiency of classical variable length codes. Hervé Jégou, Christine Guillemot |
ICASSP (4) | 2 |
| 2003 | Quantized frame expansions based on tree-structured oversampled filter banks for erasure recoveryabstractThe paper studies quantized frame expansions based on tree-structured oversampled filter banks for robust transmission of multimedia signals over erasure channels. The dependencies between the expansion coefficients introduced by the oversampled tree-structured filter banks are analysed in the z-transform domain. The capacity of correction in the presence of different erasure patterns then follows naturally. This analysis leads to the design of a reconstruction algorithm in the presence of erasures, and first without quantization noise. The impact of quantization error on the reconstruction error is then studied. For the tree-structured oversampled filter banks, we prove that reconstruction from a larger class of erasure patterns is possible as compared to regular oversampled filter banks. The reconstruction mean square error is significantly lower in the case of the tree-structured filter bank as compared to regular oversampled filter banks (Motwani, R. and Guillemot, C., 2003) and block transforms, even if the erasures are bursty in nature. Ravi Motwani, Christine Guillemot |
ICASSP (4) | 2 |
| 2003 | Subspace algorithms for error localization with DFT codesabstractWe propose two subspace algorithms for error localization with quantized DFT codes. The algorithms are similar to the MUSIC and the minimum-norm algorithms employed in array signal processing for direction of arrival (DOA) estimation. We present the algorithms in a generalized form with a variable dimension syndrome covariance matrix. Simulation results show that their localization performances are similar, and they achieve the peak values when the rank of the covariance matrix reaches the maximum value. They also perform better than the coding theoretic approach over a broad range of channel-error-to-quantization-noise ratio. Gagan Rath, Christine Guillemot |
ICASSP (4) | 2 |
| 2003 | Source multiplexed codes for error-prone channelsabstractCompression systems of real signals (images, video, audio) generate sources of information with different levels of priority, which are then encoded with variable length codes (VLC). This paper addresses the issue of robust transmission of such VLC encoded heterogeneous sources over error-prone channels. VLCs are very sensitive to channel noise: when some bits are altered, synchronization losses can occur at the receiver. This paper describes a new family of codes, called multiplexed codes, that allow to confine the de-synchronization phenomenon to low priority data while allowing to reach asymptotically the entropy bound for both (low and high priority) sources. The idea consists in creating fixed length codes for high priority information and in using the inherent redundancy to describe low priority data, hence the name "multiplexed codes". Simulation results reveal very high error resilience at almost no cost in compression efficiency. Hervé Jégou, Christine Guillemot |
ICC | 2 |
| 2003 | A subspace based method for error correction with DFT codesabstractA subspace based method for real error correction with low-pass DFT codes is proposed. The common roots of the eigen-polynomials of the syndrome covariance matrix associated with the zero eigenvalue gives the error locations. When the codevectors are quantized, the localization algorithm is modified so as to minimize the quantization noise effects. The proposed method is analogous to the subspace based spectral estimation techniques in array signal processing, but adapted to discrete error locations and only one set of syndrome coefficients. By combining the analytical results for perfect error localization with the subspace based approach, we obtain superior results than the existing approach. Gagan Rath, Christine Guillemot |
ICC | 2 |
| 2003 | Robust decoding of arithmetic codes for image transmission over error-prone channelsabstractThis paper addresses the issue of robust decoding of arithmetic codes. We first analyze dependencies between the variables involved in arithmetic coding by means of the Bayesian formalism. This provides a suitable framework for designing a soft decoding algorithm that provides high error-resilience. It also provides a natural setting for "soft synchronization", i.e., to introduce anchors favoring the likelihood of "synchronized" paths. In order to maintain the complexity of the estimation within a realistic range, a simple, yet efficient, pruning method is described. Models and algorithms are then applied to context-based arithmetic coding widely used in practical systems (e.g. JPEG-2000). Experimentation results with both theoretical sources and with real images coded with JPEG-2000 reveal very good error resilience performances. Thomas Guionnet, Christine Guillemot |
ICIP (1) | 2 |
| 2003 | 2-channel oversampled filter banks as joint source-channel codes for erasure channelsabstractThis paper studies oversampled filter banks (OFB) for robust transmission of image/video signals over erasure channels. The dependencies introduced by the OFB-based frame expansion are first analyzed in the z domain. This analysis paves the way to the design of an algorithm for message reconstruction in presence of erasures and quantization noise. Conditions for recovery from some typical erasure patterns like bursty erasures and periodic erasure patterns are derived for both orthogonal and biorthogonal filter banks. For some erasure patterns the equivalent transform coder forms a frame, hence the reconstruction could be based on operating its pseudo inverse, however at the expense of high computational cost. Alternately, the erased samples can first be reconstructed from the received ones and then signal space projection can be applied. Since in practical systems the codevectors are quantized before transmission, we study the effect of quantization noise on the reconstructed signal. The theoretical results are validated for a number of erasure patterns with image signals. Ravi Motwani, Christine Guillemot |
ICIP (1) | 2 |
| 2003 | A subspace algorithm for error localization with overcomplete frame expansions and its application to image transmissionabstractIn this paper, we propose a subspace algorithm for error localization with quantized frame expansion. The algorithm is similar to the MUSIC algorithm for direction of arrival (DOA) estimation in the array signal processing. The frames considered are a class of error-correcting frames characterized by the BCH-like property of the parity check matrices of the associated complex block codes. In particular, we consider low-pass DFT, DCT and DST frames. Simulation results show that the subspace algorithm outperforms the error locator polynomial approach in localization efficiency. For a G-M source, the DFT frame is seen to have the best performance, however, when applied to an image, the DCT and DST frames are seen to perform somewhat better than the DFT frame especially at higher channel bit error rates. Gagan Rath, Christine Guillemot |
ICIP (1) | 2 |
| 2003 | Low-rate FGS video compression based on motion-compensated spatio-temporal wavelet analysis
Jérôme Viéron, Christine Guillemot |
VCIP | 2 |
| 2003 | Real error and erasure correction with DFT codes for communication channelsabstractThis paper deals with the decoding of lowpass DFT codes in presence of both errors and erasures. First we consider the case when the codevectors are not quantized and present a coding theoretic algorithm as well as subspace based decoding algorithm. Then we consider the quantization of the codevectors and show the equivalence between the error and erasure correction, and the estimation of directions of arrival (DOA) of plane waves in array signal processing. This analogy leads to the introduction of an adaptation of the subspace based DOA estimation techniques to the problem of error and erasure correction in the real field. The adaptation incorporates a whitening operation, which is shown to be equivalent to the projection of the syndrome vector onto the orthogonal complement of the subspace spanned by the erasure locator vectors. Simulation results with a Gauss-Markov source reveal that the proposed subspace based algorithm is more efficient than the coding theoretic approach in localizing the errors. Gagan Rath, Christine Guillemot |
WCNC | 2 |
| 2003 | Soft decoding and synchronization of arithmetic codes: application to image transmission over noisy channelsabstractThis paper addresses the issue of robust and joint source-channel decoding of arithmetic codes. We first analyze dependencies between the variables involved in arithmetic coding by means of the Bayesian formalism. This provides a suitable framework for designing a soft decoding algorithm that provides high error-resilience. It also provides a natural setting for "soft synchronization", i.e., to introduce anchors favoring the likelihood of "synchronized" paths. In order to maintain the complexity of the estimation within a realistic range, a simple, yet efficient, pruning method is described. The algorithm can be placed in an iterative source-channel decoding structure, in the spirit of serial turbo codes. Models and algorithms are then applied to context-based arithmetic coding widely used in practical systems (e.g., JPEG-2000). Experimentation results with both theoretical sources and with real images coded with JPEG-2000 reveal very good error resilience performances. Thomas Guionnet, Christine Guillemot |
IEEE Trans. Image Process. | 2 |
| 2002 | Syndrome Decoding and Performance Analysis of DFT Codes with Bursty ErasuresabstractIn this paper, we analyze the performance of DFT codes with bursty erasures in the framework of syndrome decoding. Bursty erasures give rise to a syndrome decoding matrix which has very large elements depending on the code parameters and the burst length. The largeness of the elements of the syndrome decoding matrix is studied by establishing various relationships between the syndrome decoding matrix, and the parity and the generator polynomial coefficients. With a suitable model for the quantization error, the reconstruction error performance of DFT codes in the context of the syndrome decoding is analyzed and then applied to the case of bursty erasures. Simulation results with a Gauss-Markov source verify the theoretical results obtained with the assumed quantization error model. Gagan Rath, Christine Guillemot |
DCC | 2 |
| 2002 | Recursive syndrome decoding of DFT codes for bursty erasuresabstractRecently DFT codes have been considered for use as joint source-channel codes in order to provide robustness against packet loss in IP networks. In this paper, we propose a recursive syndrome decoding algorithm of DFT codes for bursty erasure patterns. The algorithm is based on the statistical properties of the quantization error and the codevector. The statistical analysis of the codevector is independent of the proposed algorithm. We also analyze the error performance of the proposed decoding scheme, and show that it can perform better than the conventional syndrome decoding of DFT codes if the message signal model and quantization error model assumptions are valid. Simulation results with a Gauss-Markov source are presented. Gagan Rath, Christine Guillemot |
ICASSP | 2 |
| 2002 | Perceptual watermarking of non i.i.d. signals based on wide spread spectrum using side informationabstractThe theoretical foundations of data hiding have been revealed by formulating the problem as message communication over a noisy channel, i.e. the host document. In light of this definition, some solutions have been proposed. Unfortunately, the performances of those methods are limited due to interference with the host signal. Considering spread spectrum information hiding with non i.i.d. Gaussian host signals and weighted distortion measures, we propose a game-theoretic resolution of the problem using side information. Gaëtan Le Guelvouit, Stéphane Pateux, Christine Guillemot |
ICIP (3) | 3 |
| 2002 | Soft decoding of multiple descriptionsabstractA multiple description scalar quantization (MDSQ) based coding system can be regarded as a source coder (quantizer) followed by a channel coder, i.e. the combination of index and codeword assignment. The redundancy, or the correlation between the descriptions, is controlled by the number of diagonals covered by the index assignment. In this paper, we analyse dependencies between the variables involved in the MDSQ coding chain and design an estimation strategy making use of part of the global model of dependencies at each time. Inference of the hidden states of the connected models is done by applying belief propagation principles on the resulting Bayesian network. We try to evidence the most appropriate form of redundancy one should introduce in the context of variable length code compressed streams in order to fight against de-synchronizations when impaired by channel noise. Thomas Guionnet, Christine Guillemot, Eric Fabre |
ICME (2) | 2 |
| 2002 | Error localization performance analysis of quantized DFT codesabstractThis paper deals with the error localization performance analysis of low-pass DFT codes in presence of quantization noise. The performance is evaluated in terms of the magnitude of the error locator polynomial computed at the error locations. It is shown that the mean squared magnitude is a function of the channel error-to-quantization noise ratio and the relative locations of the channel errors. Gagan Rath, Christine Guillemot |
ITW | 2 |
| 2001 | Application of DFT codes for robustness to erasuresabstractDiscrete Fourier transform (DFT) codes are used to provide robustness against packet losses in IP networks. First, the relationship between DFT codes and over-complete expansions using frames is established. It is shown that the message recovery by syndrome decoding of erasures and direct signal space projection from received samples are equivalent. The error performance analysis of DFT codes is used to develop packetization criteria in order to guarantee minimum reconstruction error. Gagan Rath, Xavier Henocq, Christine Guillemot |
GLOBECOM | 3 |
| 2001 | Joint source-channel turbo decoding of VLC-coded Markov sourcesabstractWe analyse the dependencies between the variables involved in the source and channel coding chain. This chain is composed of (1) a Markov source of symbols, followed by (2) a variable length source coder, and (3) a channel coder. The output process is analysed in the framework of Bayesian networks, which provide both an intuitive representation of the structure of dependencies, and a way of deriving joint (soft) decoding algorithms. Joint decoding relying on the hidden Markov model (HMM) of the global coding chain is intractable, except in trivial cases, due to the high dimensionality of the state space. We advocate instead an iterative procedure inspired from serial turbo codes, in which the three models of the coding chain are used alternately. This idea of using separately each factor of a big product model inside an iterative procedure usually requires the presence of an interleaver between successive components. We show that only one interleaver is necessary here, placed between the source coder and the channel coder. As a sub-product, we also derive a soft VLC decoder with good (and adjustable) synchronization properties. Eric Fabre, Arnaud Guyader, Christine Guillemot |
ICASSP | 3 |
| 2001 | Embedded multiple description coding for progressive image transmission over unreliable channelsabstractA multiple description scalar quantization (MDSQ) based coding system can be regarded as a source coder (quantizer) followed by a channel coder, i.e. the combination of index and codeword assignment. The redundancy, or the correlation between the descriptions, is controlled by the number of diagonals covered by the index assignment. We consider here the usage of multiple description uniform scalar quantization (that we call MDUSQ) for robust and progressive transmission of images over unreliable channels. The progressive feature is an important factor for rate control in non-stationary (varying bandwidth) communication environments. In this context, the paper describes an embedded index assignment strategy that provides improved rate-distortion performances in progressive transmission scenarios, against index assignments defined so far for MDSQ. The MDUSQ together with the embedded index assignment algorithm are incorporated into the JPEG2000 verification model. The approach is compared against a progressive multiple description scheme based on a polyphase transform (PT) decomposition of the signal. Thomas Guionnet, Christine Guillemot, Stéphane Pateux |
ICIP (1) | 2 |
| 2001 | Source adaptive TCP-compatible rate control for video over the InternetabstractThis paper describes a rate control algorithm that captures not only the behavior of TCP's congestion control avoidance mechanism but also the delay constraints of real-time streams. Building upon the TFRC (TCP-friendly rate control) protocol (see Floyd, S. et al., Proc. ACM SIGCOMM'2000, p.43-56, 2000), a new protocol based on RTP/RTCP has been designed in order to take into better account the multimedia flow characteristics in the estimation of the bandwidth prediction model parameters. Given the current channel state, the encoder and decoder buffer states and the delay constraints of the real-time video are translated into encoder rate constraints. This global rate control model, that encompasses the source buffer model as well as the end-to-end delay characteristics, allows significant reduction of source timeouts, hence minimizes the expected distortion, for a comparable usage of the TCP compatible predicted bandwidth. Jérôme Viéron, Christine Guillemot |
ICIP (1) | 2 |
| 2001 | Joint source-channel turbo decoding of entropy-coded sourcesabstractWe analyze the dependencies between the variables involved in the source and channel coding chain. This analysis is carried out in the framework of Bayesian networks, which provide both an intuitive representation for the global model of the coding chain and a way of deriving joint (soft) decoding algorithms. Three sources of dependencies are involved in the chain: (1) the source model, a Markov chain of symbols; (2) the source coder model, based on a variable length code (VLC), for example a Huffman code; and (3) the channel coder, based on a convolutional error correcting code. Joint decoding relying on the hidden Markov model (HMM) of the global coding chain is intractable, except in trivial cases. We advocate instead an iterative procedure inspired from serial turbo codes, in which the three models of the coding chain are used alternately. This idea of using separately each factor of a big product model inside an iterative procedure usually requires the presence of an interleaver between successive components. We show that only one interleaver is necessary here, placed between the source coder and the channel coder. The decoding scheme we propose can be viewed as a turbo algorithm using alternately the intersymbol correlation due to the Markov source and the redundancy introduced by the channel code. The intermediary element, the source coder model, is used as a translator of soft information from the bit clock to the symbol clock. Arnaud Guyader, Eric Fabre, Christine Guillemot, Matthias Robert |
IEEE J. Sel. Areas Commun. | 3 |
| 2001 | Indexing algorithms for Zn, An, Dn, and Dn++ lattice vector quantizersabstractThis paper describes vector indexing algorithms valid for a large class of lattices (Z/sub n/, A/sub n/, D/sub n/, and D/sub n//sup ++/ including as special cases the Gosset (E/sub 8/) and Barnes-Wall (/spl Lambda//sub 16//sup n/) lattices), widely used in audio-visual signal compression. The indexing mechanism relies on reverse lexicographic ordering of the vectors, in classes of equivalence defined as sets of vectors obtained by "signed" permutations of the components of initial vectors called leaders. The approach makes it possible to trade the size of the lookup tables-or codebooks-for arithmetic operations. The reduction of codebook sizes leads in turn to reduced encoder and decoder complexities. Pursuing the goal of best tradeoff between storage requirements and arithmetic complexity, two algorithms, based on the proposed indexing mechanisms, allowing "on-the-fly" generation of adaptive portions of codebooks are then described. These algorithms make it possible to overcome the problems of lattice truncating usually encountered in lattice vector quantization (LVQ). Combined with product codes, these indexing techniques lead to increased compression performances. Patrick Rault, Christine Guillemot |
IEEE Trans. Multim. | 2 |
| 2000 | Hybrid Sender and Receiver Driven Rate Control in Multicast Layered Video TransmissionabstractLayered coding is often proposed as a solution for rate-based congestion control of video transmission in heterogeneous environments. The problem addressed more specifically here is the design of a responsive mechanism for rate allocation in each layer, that would guarantee the best bandwidth usage for all the receivers. One key features, in contrast with existing solutions, are an optimization of the overall PSNR seen by the receivers instead of a measure of goodput. A second key feature is that the optimization does not have to be carried out on the network nodes but only by the sender. Only feedback mergers have to be supported by some nodes in the network leading to a clustering of the receivers according to their bottleneck rate of the rate they are allowed following some TCP-friendliness criteria. Fabrice Le Léannec, Jérôme Viéron, Xavier Henocq, Christine Guillemot |
ICIP | 4 |
| 2000 | Joint source and channel rate control in multicast layered video transmission
Xavier Henocq, Fabrice Le Léannec, Christine Guillemot |
VCIP | 3 |
| 1999 | Packet loss resilient MPEG-4 compliant video coding for the Internet
Fabrice Le Léannec, François Toutain, Christine Guillemot |
Signal Process. Image Commun. | 3 |
| 1998 | A Simplified Rate-Distortion Optimization Procedure Relying on Statistical Subband and Noise ModellingabstractRate-distortion optimization procedures are very often used in rate allocation problems encountered in source encoders. However, these procedures can be, depending on the number of quantizers considered, very time consuming. This paper describes a simplified approach relying on statistical signal and noise modeling. The strategy presented leads to performances comparable to a classical rate-distortion procedure with a significant reduction of computational cost, making real-time implementations very feasible. Patrick Rault, Christine Guillemot |
ICIP (2) | 2 |
| 1995 | Data-rate constrained lattice vector quantization: a new quantizing algorithm in a rate-distortion senseabstractThis paper describes an image coding scheme using lattice vector quantizers where the emphasis is put on a new lattice vector quantization approach. The adaptive signal decomposition uses wavelet packets which allow to best match the decomposition to the signal non-stationary characteristics. The quantizers used here are data rate constrained lattice vector quantizers. A classical rate-distortion algorithm [Ramchandran and Vetterli, 1993] based on two nested optimization processes allows to jointly optimize transformation and quantization and is used as a first step in the quantization procedure described here. It allows to define in each subband the lattice spacing (or scaling factor) minimizing the overall distortion for a given bit rate and to choose by a pruning algorithm the best structure of decomposition. An additional procedure developed here allows (by exploiting the lattice properties) to project a vector on a lattice point providing a better rate-distortion tradeoff. In addition, a new partitioning of the vector space allowing to divide the source of vectors into three sub-sources is introduced in order to improve the coding efficiency. In the rate distortion plane, this new quantization procedure brings a significant improvement of about 0.5-0.6 dB with respect to the classical rate-distortion algorithm. Patrice Onno, Christine Guillemot |
ICIP | 2 |
| 1994 | Cosine-Modulated Wavelets: New Results on Desgin of Arbitrary Length Filters and Optmization for Image CompressionabstractThis paper presents a procedure for designing cosine-modulated wavelet filters of arbitrary length applicable to image compression. Both criteria of frequency selectivity and of maximum coding gain are first considered in the design. In order to obtain cosine modulated bases of compactly supported M-band wavelets, additional regularity constraints are then imposed on the filters. The approach presented here allows to derive arbitrary length solutions with high stopband attenuation even in the presence of additional constraints such as regularity. The impact of the different parameters of the wavelet filters on the compression efficiency is evaluated. These filter solutions are also tested comparatively to wavelet packets in a subband coding scheme using lattice vector quantization optimized in the rate-distortion sense.> Christine Guillemot, Patrice Onno |
ICIP (1) | 1 |
| 1994 | Wavelet Packet Coding with Jointly Optimized Lattice Vector Quantization and Data Rate AllocationabstractDescribes a new coding scheme based on a joint optimization of signal decomposition and of lattice vector quantizers. The adaptive signal decomposition uses wavelet packets that allow to best match the decomposition to the signal nonstationary characteristics. In order to obtain a good tradeoff between frequency and spatial localization, nonstationary wavelet packets based on filters varying at each stage of the tree structured decomposition are considered. The complete tree is pruned into the best basis subtree which minimizes a perceptually weighted distortion for a given bit rate. Each subband is then quantized using data rate constrained lattice vector quantizers. The choice of the quantizer is usually determined by the minimization of the overall mean squared error of the reconstructed signal under the constraint of a given bit rate [Ramchandran and Vetterli, 1993]. However, this procedure is not optimum in terms of visual quality. A procedure of quantization based on the minimization of perceptually weighted distortion of each frequency band is introduced. The results show a very significant improvement both in terms of SNR and of visual quality versus the JPEG algorithm.> Patrice Onno, Christine Guillemot |
ICIP (3) | 2 |
| 1994 | Exact Reconstruction Filter Banks Using Cosine Modulation: Matrix Formalization for Arbitrary Length Prototype FiltersabstractThis paper describes generalized frequency domain matrix formalizations for two families of filter banks using different structures of polyphase component decomposition that are valid for prototype filters of arbitrary lengths and for an arbitrary number of channels M. In each case, closed form expressions of the polyphase components of the FIR prototype filter are provided and necessary and sufficient conditions on these polyphase components so that the analysis/synthesis system satisfies the perfect reconstruction condition are derived. It is shown that solutions with arbitrary lengths of the form mM-R can be obtained. For both structures, the relations between the polyphase component pairs lead naturally to fast computation algorithms, with very regular structures and low arithmetic complexity.> Christine Guillemot, Patrice Onno |
ISCAS | 1 |
| 1994 | Layered Coding Schemes for Video Transmission on ATM Networks
Christine Guillemot, Rashid Ansari |
J. Vis. Commun. Image Represent. | 1 |
| 1993 | Tradeoffs in the design of wavelet filters for image compressionabstractThis paper addresses the problem of joint optimization of wavelet transform, quantization, and data rate allocation according to mathematical criteria for high compression efficiency of image coding algorithms. The relevancy of some filter bank properties for compression purposes is evaluated. Using lattice structures, a large number of orthogonal and biorthogonal wavelet filter banks, with different properties of regularity, coding gain, phase linearity, and cross-correlation between adjacent bands are designed. Scalar and lattice vector quantization is then optimized adaptively to filter bank characteristics and to signal statistics. An appropriate choice of transition bandwidth, decreasing the energy around the Nyquist frequency without constraints of `zeros' in (omega) equals (pi) , provides by the maximum selectivity criterion filter banks close in performance to filters that we found optimum, and designed to satisfy either the maximum coding gain or minimum cross-correlation criterion. For a lower transition bandwidth, the increased regularity has for effect to increase the coding gain, to reach a maximum coding gain for maximally regular Daubechies filters. When comparing results of coding with the optimal orthogonal wavelet filter bank with those provided by a maximally frequency selective biorthogonal solution with same regularity it is observed that for a comparable peak SNR the contours are better reconstructed with biorthogonal solutions. Patrice Onno, Christine Guillemot |
VCIP | 2 |
| 1991 | M-channel nonrectangular wavelet representation for 2-D signals: basis for quincunx sampled signalsabstractThe authors have described the framework underlying continuous and discrete families of nonseparable two-dimensional wavelets and an M-channel nonrectangular multiresolution wavelet representation for L/sup 2/(R/sup 2/) functions and I/sup 2/(Z/sup 2/) sequences. Focusing on the case of quincunx sampled signals, solutions of digital filter banks for implementing the decomposition with different characteristics of linear phase and regularity for smoothing and wavelet functions are provided.> Christine Guillemot, A. Enis Çetin, Rashid Ansari |
ICASSP | 1 |
| 1990 | Polynomial transform computation of the 2-D DCTabstractA 2-D DCT (discrete cosine transform) algorithm based on a direct polynomial approach is presented. The resulting algorithm reduces the number of both multiplications and additions compared to previous algorithms. It is shown that, although being mathematically involved, it possesses a clean, butterfly-based structure. Tables comparing the number of operations are provided, as well as flowgraphs.> Pierre Duhamel, Christine Guillemot |
ICASSP | 2 |
| 1989 | Trade-off's in the computation of mono- and multi-dimensional DCT'sabstractAn overview of some alternative algorithms for one- and two-dimensional DCTs (discrete cosine transforms) is given. Operation counts are derived for typical examples useful in image processing. It is shown that it is possible to generalize the 2-D schemes to 3-D DCTs as well. The result is that a 3-D DCT can be obtained from a 3-D DFT (discrete Fourier transform) of the same size on reals at the cost of permutations and O(3/2N/sup 3/) multiplications. The scheme involves rotations on eight output points at a time. Improvements through scaling are discussed, and implementation issues (both in hardware and software) are addressed.> Martin Vetterli, Pierre Duhamel, Christine Guillemot |
ICASSP | 3 |