VLDB 2026 Research / reviewers in the wild / expert
Thierry Dumas
dblp:186/3059
· DBLP profile ↗
12ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-7751-5219ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Partition Tree Search Acceleration for VVC: Survey and Evaluation with VTM EvolutionabstractVVC achieves up to 50% bit-rate savings over HEVC at the cost of increased encoding complexity, largely due to the QTMTT partitioning structure. This work surveys the evolution of the VVC Test Model (VTM) and evaluates partitioning acceleration techniques considering changes in complexity and internal heuristics across VTM versions. M. E. A. Kherchouche, Franck Galpin, Thierry Dumas, Daniel Ménard |
DCC | 3 |
| 2025 | Complexity Reduction Study Based on RD Costs Approximation for VVC Intra PartitioningabstractThis paper presents a comparison study of two machine-learning techniques to accelerate the Versatile Video Coding (VVC) intra-partitioning process in QuadTree (QT) configuration: a regression model for predicting Rate-Distortion (RD) costs and a Deep Q-Network (DQN)-based Reinforcement Learning (RL) approach that models partitioning as a Markov Decision Process (MDP). Both methods are size-independent and utilize neighboring RD costs and threshold values to optimize the splits of Coding Units (CUs). M. E. A. Kherchouche, Franck Galpin, Thierry Dumas, F. Schnitzler, Daniel Ménard |
DCC | 3 |
| 2025 | Enhanced Template-based Intra Mode Derivation with Adaptive Block Vector Replacement
Jiaqi Zhang 0007, Jiaye Fu, Chuanmin Jia, Siwei Ma 0001, Karam Naser, Thierry Dumas, Saurabh Puri, Milos Radosavljevic |
PCS | 6 |
| 2025 | Advanced Neural Network-Based Video Coding Technologies for Intra Prediction and In-Loop FilteringabstractThe past decade has witnessed the huge success of deep learning in well-known artificial intelligence applications such as face recognition, autonomous driving, and large language model like ChatGPT. Recently, the application of deep learning has been extended to a much wider range, with Neural Network-Based Video Coding (NNVC) being one of them. NNVC can be performed at two different levels: embedding neural network-based (NN-based) coding tools into a classical video compression framework or building the entire compression framework upon neural networks. This article elaborates our studies in response to the recent exploration efforts in JVET (Joint Video Experts Team of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC29) in the name of NNVC, falling in the former category. Specifically, in this article, we propose two advanced NN-based video coding technologies, i.e., NN-based intra prediction and NN-based in-loop filtering, which have been investigated for several meeting cycles in JVET and then adopted into the reference software, i.e., NNVC. In addition, we further propose a Small Ad-hoc Deep-Learning Library (SADL), which provides integer-based inference capabilities for neural networks to ensure interoperability across different systems. SADL has been adopted as the inference platform of all neural networks in NNVC. Extensive experiments on top of the NNVC have been conducted to evaluate the effectiveness of the proposed techniques. Compared with VTM-11.0_nnvc, the proposed two NN-based coding tools jointly achieve {11.94%, 21.86%, 22.59%}, {9.18%, 19.76%, 20.92%}, and {10.63%, 21.56%, 23.02%} BD-rate reductions on average for {Y, Cb, Cr} under random-access, low-delay, and all-intra configurations, respectively. Yue Li 0015, Chaoyi Lin, Kai Zhang 0007, Li Zhang 0006, Franck Galpin, Thierry Dumas, Muhammed Coban, Jacob Ström, Du Liu, Kenneth Andersson |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2024 | RD-cost Regression Speed Up Technique for VVC Intra Block PartitioningabstractThe last standard Versatile Video Codec (VVC) aims to improve the compression efficiency by saving around 50% of bitrate at the same quality compared to its predecessor High Efficiency Video Codec (HEVC). However, this comes with higher encoding complexity mainly due to a much larger number of block splits to be tested on the encoder side. This paper proposes an acceleration of the VVC partitioning based on a multi-output regression model that predicts a suitable split mode for 32 × 32 Coding Unit (CU). Experimental results show that our approach improves complexity trade-offs flexibility while adding better complexity trade-offs points compared to the original encoder. M. E. A. Kherchouche, Franck Galpin, Thierry Dumas, Daniel Ménard |
ICASSP | 3 |
| 2021 | Revisiting the Sample Adaptive Offset post-filter of VVC with Neural-NetworksabstractThe Sample Adaptive Offset (SAO) filter has been introduced in HEVC to reduce general coding and banding artefacts in the reconstructed pictures, in complement to the De-Blocking Filter (DBF) which reduces artifacts at block boundaries specifically. The new video compression standard Versatile Video Coding (VVC) reduces the BD-rate by about 36% at the same reconstruction quality compared to HEVC. It implements an additional new in-loop Adaptive Loop Filter (ALF) on top of the DBF and the SAO filter, the latter remaining unchanged compared to HEVC. However, the relative performance of SAO in VVC has been lowered significantly. In this paper, it is proposed to revisit the SAO filter using Neural Networks (NN). The general principles of the SAO are kept, but the a-priori classification of SAO is replaced with a set of neural networks that determine which reconstructed samples should be corrected and in which proportion. Similarly to the original SAO, some parameters are determined at the encoder side and encoded per CTU. The average BD-rate gain of the proposed SAO improves VVC by at least 2.3% in Random Access while the overall complexity is kept relatively small compared to other NN-based methods. Philippe Bordes, Franck Galpin, Thierry Dumas, Pavel Nikitin |
PCS | 3 |
| 2021 | Combined Neural Network-based Intra Prediction and Transform SelectionabstractThe interactions between different tools added successively to a block-based video codec are critical to its ratedistortion efficiency. In particular, when deep neural network-based intra prediction modes are inserted into a block-based video codec, as the neural network-based prediction function cannot be easily characterized, the adaptation of the transform selection process to the new modes can hardly be performed manually. That is why this paper presents a combined neural network-based intra prediction and transform selection for a block-based video codec. When putting a single neural network-based intra prediction mode and the learned prediction of the selected LFNST pair index into VTM-8.0, -3.71%, -3.17%, and -3.37% of mean BD-rate reduction in all-intra is obtained. Thierry Dumas, Franck Galpin, Philippe Bordes |
PCS | 1 |
| 2021 | Neural Network based Inter bi-prediction BlendingabstractThis paper presents a learning-based method to improve bi-prediction in video coding. In conventional video coding solutions, the motion compensation of blocks from already decoded reference pictures stands out as the principal tool used to predict the current frame. Especially, the bi-prediction, in which a block is obtained by averaging two different motion-compensated prediction blocks, significantly improves the final temporal prediction accuracy. In this context, we introduce a simple neural network that further improves the blending operation. A complexity balance, both in terms of network size and encoder mode selection, is carried out. Extensive tests on top of the recently standardized VVC codec are performed and show a BD-rate improvement of −1.4% in random access configuration for a network size of fewer than 10k parameters. We also propose a simple CPU-based implementation and direct network quantization to assess the complexity/gains tradeoff in a conventional codec framework. Franck Galpin, Philippe Bordes, Thierry Dumas, Pavel Nikitin, Fabrice Le Léannec |
VCIP | 3 |
| 2021 | Iterative Training of Neural Networks for Intra PredictionabstractThis paper presents an iterative training of neural networks for intra prediction in a block-based image and video codec. First, the neural networks are trained on blocks arising from the codec partitioning of images, each paired with its context. Then, iteratively, blocks are collected from the partitioning of images via the codec including the neural networks trained at the previous iteration, each paired with its context, and the neural networks are retrained on the new pairs. Thanks to this training, the neural networks can learn intra prediction functions that both stand out from those already in the initial codec and boost the codec in terms of rate-distortion. Moreover, the iterative process allows the design of training data cleansings essential for the neural network training. When the iteratively trained neural networks are put into H.265 (HM-16.15), -4.2% of mean BD-rate reduction is obtained, i.e. -1.8% above the state-of-the-art. By moving them into H.266 (VTM-5.0), the mean BD-rate reduction reaches -1.9%. Thierry Dumas, Franck Galpin, Philippe Bordes |
IEEE Trans. Image Process. | 1 |
| 2020 | Context-Adaptive Neural Network-Based Prediction for Image CompressionabstractThis paper describes a set of neural network architectures, called Prediction Neural Networks Set (PNNS), based on both fully-connected and convolutional neural networks, for intra image prediction. The choice of neural network for predicting a given image block depends on the block size, hence does not need to be signalled to the decoder. It is shown that, while fully-connected neural networks give good performance for small block sizes, convolutional neural networks provide better predictions in large blocks with complex textures. Thanks to the use of masks of random sizes during training, the neural networks of PNNS well adapt to the available context that may vary, depending on the position of the image block to be predicted. When integrating PNNS into a H.265 codec, PSNRrate performance gains going from 1:46% to 5:20% are obtained. These gains are on average 0:99% larger than those of prior neural network based methods. Unlike the H.265 intra prediction modes, which are each specialized in predicting a specific texture, the proposed PNNS can model a large set of complex textures. Thierry Dumas, Aline Roumy, Christine Guillemot |
IEEE Trans. Image Process. | 1 |
| 2018 | Autoencoder Based Image Compression: Can the Learning be Quantization Independent?abstractThis paper explores the problem of learning transforms for image compression via autoencoders. Usually, the rate-distortion performances of image compression are tuned by varying the quantization step size. In the case of autoencoders, this in principle would require learning one transform per rate-distortion point at a given quantization step size. Here, we show that comparable performances can be obtained with a unique learned transform. The different rate-distortion points are then reached by varying the quantization step size at test time. This approach saves a lot of training time. Thierry Dumas, Aline Roumy, Christine Guillemot |
ICASSP | 1 |
| 2017 | Image compression with Stochastic Winner-Take-All Auto-EncoderabstractThis paper addresses the problem of image compression using sparse representations. We propose a variant of autoencoder called Stochastic Winner-Take-All Auto-Encoder (SWTA AE). “Winner-Take-All” means that image patches compete with one another when computing their sparse representation and “Stochastic” indicates that a stochastic hyperparameter rules this competition during training. Unlike auto-encoders, SWTA AE performs variable rate image compression for images of any size after a single training, which is fundamental for compression. For comparison, we also propose a variant of Orthogonal Matching Pursuit (OMP) called Winner-Take-All Orthogonal Matching Pursuit (WTA OMP). In terms of rate-distortion trade-off, SWTA AE outperforms auto-encoders but it is worse than WTA OMP. Besides, SWTA AE can compete with JPEG in terms of rate-distortion. Thierry Dumas, Aline Roumy, Christine Guillemot |
ICASSP | 1 |