Félix Henry

dblp:23/4897 · DBLP profile ↗
← Back
22ranked-venue papers
1as first author
13since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Upsampling Improvement for Overfitted Neural Coding
abstract
Neural image compression, based on auto-encoders and overfitted representations, relies on a latent representation of the coded signal. This representation needs to be compact and uses low resolution feature maps. In the decoding process, those latents are upsampled and filtered using stacks of convolution filters and non linear elements to recover the decoded image.Therefore, the upsampling process is crucial in the design of a neural coding scheme and is of particular importance for overfitted codecs where the network parameters, including the upsampling filters, are part of the representation.This paper addresses the improvement of the upsampling process in order to reduce its complexity and limit the number of parameters. A new upsampling structure is presented whose improvements are illustrated within the Cool-Chic overfitted image coding framework. The proposed approach offers a rate reduction of 4.7%. The source code is available [1].
Pierrick Philippe, Théo Ladune, Gordon Clare, Félix Henry, Théophile Blard, Thomas Leguay
ISCAS4
2025 Efficient Sub-pixel Motion Compensation in Learned Video Codecs
Théo Ladune, Thomas Leguay, Pierrick Philippe, Gordon Clare, Félix Henry
PCS5
2024 Cool-Chic: Perceptually Tuned Low Complexity Overfitted Image Coder
abstract
This paper summarises the design of the Cool-Chic candidate for the Challenge on Learned Image Compression. This candidate attempts to demonstrate that neural coding methods can lead to low complexity and lightweight image decoders while still offering competitive performance. The approach is based on the already published overfitted lightweight neural networks Cool-Chic, further adapted to the human subjective viewing targeted in this challenge.
Théo Ladune, Pierrick Philippe, Gordon Clare, Félix Henry, Thomas Leguay
DCC4
2023 GOP-Based Latent Refinement for Learned Video Coding
abstract
This paper presents a method allowing learned video encoders to apply arbitrary latent refinement strategies to serve as RateDistortion Optimization (RDO) at the time of encoding. To do so, a latent domain search is applied on an initial latent representation of the video signal. This search is implemented as a set of iterations, each of which performs a gradient descent with back-propagation of error defined by a Lagrangian RD cost. This cost function is intentionally chosen to be the same as the cost function that was used during the end-to-end model training, except that instead of updating model weights, each iteration fine-tunes the latent representation itself. Moreover, a temporal look-ahead is integrated in the cost function of I and P frames to take into account the cascade effect of their latent fine-tuning on subsequent frames in the Group of Pictures (GOP). The experiments show that the proposed latent space RDO method can improve by 11.6% and 9.4% in terms of BD-BR coding efficiency in Random-Access (RA) and All-Intra (AI) configurations, when applied on top a high-performance opensource end-to-end codec.
Mohsen Abdoli, Gordon Clare, Félix Henry
ICASSP3
2023 COOL-CHIC: Coordinate-based Low Complexity Hierarchical Image Codec
abstract
We introduce COOL-CHIC, a Coordinate-based Low Complexity Hierarchical Image Codec. It is a learned alternative to autoencoders with 629 parameters and 680 multiplications per decoded pixel. COOL-CHIC offers compression performance close to modern conventional MPEG codecs such as HEVC and is competitive with popular autoencoder-based systems. This method is inspired by Coordinate-based Neural Representations, where an image is represented as a learned function which maps pixel coordinates to RGB values. The parameters of the mapping function are then sent using entropy coding. At the receiver side, the compressed image is obtained by evaluating the mapping function for all pixel coordinates. COOL-CHIC implementation is made open-source1.
Théo Ladune, Pierrick Philippe, Félix Henry, Gordon Clare, Thomas Leguay
ICCV3
2023 Low-Complexity Overfitted Neural Image Codec
abstract
We propose a neural image codec at reduced complexity which overfits the decoder parameters to each input image. While autoencoders perform up to a million multiplications per decoded pixel, the proposed approach only requires 2300 multiplications per pixel. Albeit low-complexity, the method rivals autoencoder performance and surpasses HEVC performance under various coding conditions. Additional lightweight modules and an improved training process provide a 14% rate reduction with respect to previous overfitted codecs, while offering a similar complexity. This work is made open-source at http://orange-opensource.github.io/Cool-Chic/.
Thomas Leguay, Théo Ladune, Pierrick Philippe, Gordon Clare, Félix Henry, Olivier Déforges
MMSP5
2022 Motion Compensation-based Low-Complexity Decoder Side Depth Estimation for MPEG Immersive Video
abstract
Decoder-Side Depth Estimation (DSDE) is a system firstly enabled in the novel MPEG Immersive Video (MIV) coding standard. In DSDE, only texture components are coded, while the depth is estimated at the decoder-side. This is motivated by previous work, which has shown high coding gain and pixel rate savings in DSDE. However, the computational complexity remains a concern, as high quality depth search has a high runtime and memory requirement. In this work we extend the concept of depth estimation to depth recovery. Using this mode, the decoder-side depth information is recovered through motion compensation utilizing the displacement vectors contained in the texture bitstream. This strategy enables us to replace most of the complex depth estimation processes with a simple motion compensation step, a decision that is drawn on the encoder-side and signaled per coding unit. With only minor losses in terms of synthesis PSNR and similar perceptual quality in terms of MS-SSIM, the complexity is significantly reduced. Depending on the acceptable loss, up to 80 % of the moving objects depth may be motion compensated instead of estimated by a depth estimator translating into a speed-up of a factor of 104 for inter-frames compared to the reference depth estimator.
Patrick Garus, Félix Henry, Thomas Maugey, Christine Guillemot
MMSP2
2022 Decoder Side Multiplane Images using Geometry Assistance SEI for MPEG Immersive Video
abstract
The MPEG Immersive Video (MIV) standard enables a novel technology denoted as decoder side depth estimation (DSDE) by introducing a dedicated Geometry Absent profile. In DSDE only texture information is coded and the corresponding geometry is reconstructed on the decoder side. MIV further enables the coding of side-information useful to the geometry reconstruction, denoted as Geometry Assistance SEI message. An emerging format for immersive video are Multiplane Images, which is investigated for feasibility in coding systems due to their promising rendering quality with complex sequences. In this work, we show that MIV can be used to construct block-based Multiplane Images on the decoder-side and to enhance the view synthesis performance utilizing the Geometry Assistance SEI. In a complexity-aware setting using only 32 planes, up to 6 dB of quality improvement is achieved compared to the reference.
Patrick Garus, Félix Henry, Thomas Maugey, Christine Guillemot
MMSP2
2022 A Study of Conventional and Learning-Based Depth Estimators for Immersive Video Transmission
abstract
Obtaining an accurate depth map of a scene is very important for major applications like immersive video, robotics, autonomous driving, and many more. The different methods to estimate depths can be classified as conventional and learning-based methods. While these methods have been studied for their depth accuracy, less attention has been paid to studying their performance in the use case of depth image-based rendering (DIBR). Here we study and evaluate two conventional methods and five learning-based methods for a real-world use case of immersive video transmission in the context of MPEG-I. The user-requested views are synthesized using Test Model for Immersive Video (TMIV) from the depth maps obtained by all methods and original texture views. The synthesized images are compared with their original counterparts using various quality metrics.
Smitha Lingadahalli Ravi, Marta Milovanovic, Luce Morin, Félix Henry
MMSP4
2022 Depth Patch Selection for Decoder-Side Depth Estimation in MPEG Immersive Video
abstract
The MPEG immersive video (MIV) standard has been developed to efficiently compress volumetric video content and enable an immersive user experience. MIV deals with an enormous amount of data that comes in the form of multi-view plus depth videos, which is efficiently reduced in the process of pruning, by tackling the redundancies among the views. This paper presents a novel approach for improving the existing immersive video coding scheme. The proposed approach reduces the amount of transmitted depth data, leveraging the fact that the depth information is partially contained in texture videos. The study proposes a method that ensures a reliable recovery of depths at the decoder-side. This method provides BD-rate improvements on both high and low bitrate ranges, with up to 22.57% Y-PSNR, 25.76%VMAF, 24.07% MS-SSIM, and 22.94% IV-PSNR metric gain, given a low bitrate setting.
Marta Milovanovic, Félix Henry, Marco Cagnazzo
PCS2
2022 Immersive Video Coding: Should Geometry Information Be Transmitted as Depth Maps?
abstract
Immersive video often refers to multiple views with texture and scene geometry information, from which different viewports can be synthesized on the client side. To design efficient immersive video coding solutions, it is desirable to minimize bitrate, pixel rate and complexity. We investigate whether the classical approach of sending the geometry of a scene as depth maps is appropriate to serve this purpose. Previous work shows that bypassing depth transmission entirely and estimating depth at the client side improves the synthesis performance while saving bitrate and pixel rate. In order to understand if the encoder side depth maps contain information that is beneficial to be transmitted, we first explore a hybrid approach which enables partial depth map transmission using a block-based RD-based decision in the depth coding process. This approach reveals that partial depth map transmission may improve the rendering performance but does not present a good compromise in terms of compression efficiency. This led us to address the remaining drawbacks of decoder side depth estimation: complexity and depth map inaccuracy. We propose a novel system that takes advantage of high quality depth maps at the server side by encoding them into lightweight features that support the depth estimator at the client side. These features allow reducing the amount of data that has to be handled during decoder side depth estimation by 88%, which significantly speeds up the cost computation and the energy minimization of the depth estimator. Furthermore, −46.0% and −37.9% average synthesis BD-Rate gains are achieved compared to the classical approach with depth maps estimated at the encoder.
Patrick Garus, Félix Henry, Joël Jung, Thomas Maugey, Christine Guillemot
IEEE Trans. Circuits Syst. Video Technol.2
2021 Patch Decoder-Side Depth Estimation In Mpeg Immersive Video
abstract
This paper presents a new approach for achieving bitrate and pixel rate reduction in the MPEG immersive video coding setting. We demonstrate that it is possible to avoid the transmission of some depth information in the Test Model for Immersive Video (TMIV) by estimating it at the receiver's side. Although the transmitted information in TMIV is considered as non-redundant, we show that it is possible to improve this algorithm. This method provides 3.4%, 9.0%, and 12.1% average BD-rate gain for natural content on high, medium, and low bitrate, respectively, with up to respectively 12.3%, 16.0%, and 18.4% peak reductions. Moreover, it preserves the perceptual quality as measured with MS-SSIM and VMAF metrics. Additionally, it decreases the pixel rate by 8.3% for each test sequence.
Marta Milovanovic, Félix Henry, Marco Cagnazzo, Joël Jung
ICASSP2
2021 Overview of the Screen Content Support in VVC: Applications, Coding Tools, and Performance
abstract
In an increasingly connected world, consumer video experiences have diversified away from traditional broadcast video into new applications with increased use of non-camera-captured content such as computer screen desktop recordings or animations created by computer rendering, collectively referred to as screen content. There has also been increased use of graphics and character content that is rendered and mixed or overlaid together with camera-generated content. The emerging Versatile Video Coding (VVC) standard, in its first version, addresses this market change by the specification of low-level coding tools suitable for screen content. This is in contrast to its predecessor, the High Efficiency Video Coding (HEVC) standard, where highly efficient screen content support is only available in extension profiles of its version 4. This paper describes the screen content support and the five main low-level screen content coding tools in VVC: transform skip residual coding (TSRC), block-based differential pulse-code modulation (BDPCM), intra block copy (IBC), adaptive color transform (ACT), and the palette mode. The specification of these coding tools in the first version of VVC enables the VVC reference software implementation (VTM) to achieve average bit-rate savings of about 41% to 61% relative to the HEVC test model (HM) reference software implementation using the Main 10 profile for 4:2:0 screen content test sequences. Compared to the HM using the Screen-Extended Main 10 profile and the same 4:2:0 test sequences, the VTM provides about 19% to 25% bit-rate savings. The same comparison with 4:4:4 test sequences revealed bit-rate savings of about 13% to 27% for$Y'C_{B}C_{R}$and of about 6% to 14% for$R'G'B'$screen content. Relative to the HM without the HEVC version 4 screen content coding extensions, the bit-rate savings for 4:4:4 test sequences are about 33% to 64% for$Y'C_{B}C_{R}$and 43% to 66% for$R'G'B'$screen content.
Tung Nguyen 0001, Xiaozhong Xu, Félix Henry, Ru-Ling Liao, Mohammed Golam Sarwer, Marta Karczewicz, Yung Hsuan Chao, Jizheng Xu, Shan Liu 0001, Detlev Marpe, Gary J. Sullivan
IEEE Trans. Circuits Syst. Video Technol.3
2019 Transform Coefficient Coding for Screen Content in Versatile Video Coding (VVC)
abstract
A transform coefficient coding scheme is proposed for 4 × 4 blocks in Versatile Video Coding (VVC), targeting screen content applications. The proposed algorithm, called Unary Bitplane Coding (UBC), uses unary codes of the coefficient amplitudes and represents each block by their bitplanes. This representation allows exploiting further contextual information for source separation during the entropy coding. Experiments in the Joint Exploration test Model (JEM) show that replacing the existing transform coding with UBC only for 4 × 4 blocks brings on average 2.8% and 3.4% BD-R gain in the random access and all intra modes, respectively.
Mohsen Abdoli, Félix Henry, Patrice Brault, Frédéric Dufaux, Pierre Duhamel
ICASSP2
2019 Intra Block-DPCM with Layer Separation of Screen Content in VVC
abstract
An intra coding algorithm with layer separation is proposed. This algorithm is designed on top of an adopted tool in VVC, called Block DPCM (BDPCM), and benefits from texture information in a neighborhood to derive intensity levels of background and foreground layers. This information is used to reduce large rate of residual in case of incorrect layer prediction by BDPCM. For this purpose, three inter-layer transition states are defined that are either implicitly or explicitly conveyed to the decoder. Once a transition is signaled, the decoder corrects the prediction value using the derived layer information. Experiments on screen contents show a BD-rate gain of about 10% percent over VVC Test Model (VTM) and 1% over the regular BDPCM, with the cost of computational complexity.
Mohsen Abdoli, Félix Henry, Patrice Brault, Frédéric Dufaux, Pierre Duhamel, Pierrick Philippe
ICIP2
2018 Short-Distance Intra Prediction of Screen Content in Versatile Video Coding (VVC)
abstract
A novel intra prediction algorithm is proposed to improve the coding performance of screen content for the emerging Versatile Video Coding (VVC) standard. The algorithm, called in-loop residual coding with scalar quantization, employs in-block pixels as reference rather than the regular out-block ones. To this end, an additional in-loop residual signal is used to partially reconstruct the block at the pixel level, during the prediction. The proposed algorithm is essentially designed to target high detail textures, where deep block partitioning structure is required. Therefore, it is implemented to operate on 4× 4 blocks only, where further block split is not allowed and the standard algorithm is still unable to properly predict the texture. Experiments in the Joint Exploration Model (JEM) reference software show that the proposed algorithm brings a Bjontegaard Delta (BD)-rate gain of 13% on synthetic content, with a negligible computational complexity overhead at both encoder and decoder sides.
Mohsen Abdoli, Félix Henry, Patrice Brault, Pierre Duhamel, Frédéric Dufaux
IEEE Signal Process. Lett.2
2017 Intra prediction using in-loop residual coding for the post-HEVC standard
abstract
A few years after standardization of the High Efficiency Video Coding (HEVC), now the Joint Video Exploration Team (JVET) group is exploring post-HEVC video compression technologies. In the intra prediction domain, this effort has resulted in an algorithm with 67 internal modes, new filters and tools which significantly improve HEVC. However, the improved algorithm still suffers from the long distance prediction inaccuracy problem. In this paper, we propose an In-Loop Residual coding Intra Prediction (ILR-IP) algorithm which utilizes inner-block reconstructed pixels as references to reduce the distance from predicted pixels. This is done by using the ILR signal for partially reconstructing each pixel, right after its prediction and before its block-level out-loop residual calculation. The ILR signal is decided in the rate-distortion sense, by a brute-force search on a QP-dependent finite codebook that is known to the decoder. Experiments show that the proposed ILR-IP algorithm improves the existing method in the Joint Exploration Model (JEM) up to 0.45% in terms of bit rate saving, without complexity overhead at the decoder side.
Mohsen Abdoli, Félix Henry, Patrice Brault, Pierre Duhamel, Frédéric Dufaux
MMSP2
2015 Mode Dependent Vector Quantization with a rate-distortion optimized codebook for residue coding in video compression
abstract
The High Efficiency Video Coding standard (HEVC) supports a total of 35 intra prediction modes which aim at reducing spatial redundancy by exploiting pixel correlation within a local neighborhood. In this paper, we show that spatial correlation remains after intra prediction, leading to high energy prediction residues. We propose a novel scheme for encoding the prediction residues using a Mode Dependent Vector Quantization (MDVQ) which aims at reducing the redundancy in residual domain. The MDVQ codebook is optimized in a rate-distortion (RD) sense. Experimental results show that the codebook can be independent of the quantization parameter (QP) with no loss in terms of coding efficiency. A bitrate reduction of 1.1% on average compared to HEVC can be achieved, while further tests indicate that codebook adaptivity could substantially improve the performance.
Bihong Huang, Félix Henry, Christine Guillemot, Philippe Salembier
ICASSP2
2012 Multiple sign bits hiding for High Efficiency Video Coding
abstract
High Efficiency Video Coding (HEVC) is the next-generation video coding standard currently under development, which has demonstrated substantial bit savings (rate reduction by approximately half) compared to H.264/AVC. This paper presents the multiple sign bits hiding scheme that was adopted into the committee draft of HEVC at the 8th JCT-VC meeting. In HEVC, the quantized transform coefficients are entropy-coded in groups of 16 coefficients for each transform unit. With multiple sign bits hiding, for coefficient groups that satisfy certain conditions, the sign of the first non-zero coefficient along the scanning path is not explicitly transmitted in the bitstream and instead is inferred from the parity of the sum of all non-zero coefficients in that coefficient group at the decoder. To ensure the matching between the hidden sign and the parity of the sum of all non-zero coefficients, a parity adjustment method is employed at the encoder based on rate-distortion optimization or distortion minimization. Compared with conventional video coding schemes where quantization and coefficient coding are separately designed, the multiple sign bits hiding scheme in HEVC represents a joint quantization and coefficient coding design and provides consistent rate-distortion performance gains for all standard test sequences under standard test conditions.
Xiang Yu 0001, Dake He, Félix Henry, Gordon Clare
VCIP4
2012 Parallel Scalability and Efficiency of HEVC Parallelization Approaches
abstract
Unlike H.264/advanced video coding, where parallelism was an afterthought, High Efficiency Video Coding currently contains several proposals aimed at making it more parallel-friendly. A performance comparison of the different proposals, however, has not yet been performed. In this paper, we will fill this gap by presenting efficient implementations of the most promising parallelization proposals, namely tiles and wavefront parallel processing (WPP). In addition, we present a novel approach called overlapped wavefront (OWF), which achieves higher performance and efficiency than tiles and WPP. Experiments conducted on a 12-core system running at 3.33 GHz show that our implementations achieve average speedups, for 4k sequences, of 8.7, 9.3, and 10.7 for WPP, tiles, and OWF, respectively.
Chi Ching Chi, Mauricio Alvarez-Mesa, Ben H. H. Juurlink, Gordon Clare, Félix Henry, Stéphane Pateux, Thomas Schierl
IEEE Trans. Circuits Syst. Video Technol.5
2012 Transform Coefficient Coding in HEVC
abstract
This paper describes transform coefficient coding in the draft international standard of High Efficiency Video Coding (HEVC) specification and the driving motivations behind its design. Transform coefficient coding in HEVC encompasses the scanning patterns and coding methods for the last significant coefficient, significance map, coefficient levels, and sign data. Special attention is paid to the new methods of last significant coefficient coding, multilevel significance maps, high-throughput binarization, and sign data hiding. Experimental results are provided to evaluate the performance of transform coefficient coding in HEVC.
Joel Sole, Rajan L. Joshi, Tianying Ji, Marta Karczewicz, Gordon Clare, Félix Henry, Alberto Duenas
IEEE Trans. Circuits Syst. Video Technol.7
1999 Rate-Distorsion Efficiency of Zerotree Coders
abstract
Although zerotree coders are extremely simple in terms of algorithmic complexity and structure, their rate-distortion tradeoff is among the best known in the literature. This paper intends to provide an explanation to this phenomenon, by explicitly showing which part of the algorithm is at its origin. This is somewhat proved by introducing the same mechanism in other type of algorithms, which in turn show a similar rate/distortion performance. More specifically, we compare three coders: The first one is the original zerotree coder, the other ones are two "zerotree-like" coders, where coefficient significance is determined by thresholding (first coder) and by by rate-distortion optimization (second coder). These coders exhibit similar behaviour and performances. Thresholding seems to corresponds to a rate-distortion optimal way of isolating significant data. We show the impact that this procedure has on the distribution of the quantization noise. In particular, thresholding removes the coefficients responsible for the non-uniformity of the quantization noise. The significance map is encoded using classical lossless techniques. All three coders are shown to have (almost) equal rate-distortion characteristics.
Félix Henry, Pierre Duhamel
ICIP (1)1