Michael Schäfer 0003

dblp:67/6156-3 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0003-0309-3161ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Nonlinear Modifications of Transform Coefficients in VVC Intra Coding
abstract
With the emergence of the Versatile Video Coding standard (VVC), novel coding tools like matrix-based intra prediction and low-frequency non-separable transforms have been developed based on data-driven optimization methods. Motivated by the growing compression efficiency of learned nonlinear transforms in image coding, we incorporate a neural-network-based coding tool into the transform coding stage of VVC. First, a nonlinear update of the transform coefficients is applied which aims at decreasing the expected bitrate cost. Then, on the decoder side, a filter is applied before the synthesis transform to improve the reconstruction quality. These networks are jointly trained and have been added to the rate-distortion optimized quantization process. The integration of these approaches into the VTM software leads to bitrate savings between 0.94% and 2.12% in terms of the Bjøntegaard-Delta rate. We furthermore conducted multiple experiments on reducing the memory complexity of the networks by exploiting structural similarities of the intra prediction modes and block symmetries. As a result, we demonstrate that the number of distinct networks can be reduced without significantly diminishing the coding gain of the tool. With memory usage reduced by 90%, we achieve bitrate savings between 1.07% and 1.94%.
Michael Schäfer 0003, Florian Borzechowski, Heiko Schwarz, Jonathan Pfaff, Detlev Marpe, Thomas Wiegand 0001
ICIP1
2024 Optimizing Learned Image Compression On Scalar and Entropy-Constraint Quantization
abstract
The continuous improvements on image compression with variational autoencoders have lead to learned codecs competitive with conventional approaches in terms of rate-distortion efficiency. Nonetheless, taking the quantization into account during the training process remains a problem, since it produces zero derivatives almost everywhere and needs to be replaced with a differentiable approximation which allows end-to-end optimization. Though there are different methods for approximating the quantization, none of them model the quantization noise correctly and thus, result in suboptimal networks. Hence, we propose an additional finetuning training step: After conventional end-to-end training, parts of the network are retrained on quantized latents obtained at the inference stage. For entropy-constraint quantizers like Trellis-Coded Quantization, the impact of the quantizer is particularly difficult to approximate by rounding or adding noise as the quantized latents are interdependently chosen through a trellis search based on both the entropy model and a distortion measure. We show that retraining on correctly quantized data consistently yields additional coding gain for both uniform scalar and especially for entropy-constraint quantization, without increasing inference complexity. For the Kodak test set, we obtain average savings between $1 \%$ and $2 \%$, and for the TecNick test set up to $2.2 \%$ in terms of Bjøntegaard-Delta bitrate.
Florian Borzechowski, Michael Schäfer 0003, Heiko Schwarz, Jonathan Pfaff, Detlev Marpe, Thomas Wiegand 0001
ICIP2
2024 Nonlinear Transform Coding for VVC Intra Coding
abstract
Modern hybrid video codecs like Versatile Video Coding (VVC) heavily rely on transform coding tools. Given a prediction signal at the encoder, the residual is transformed using trigonometric transforms. Rate-distortion-optimized quantization (RDOQ) and entropy coding of the transformed residual is well-understood due to the orthogonality and the energy compaction of these transforms. Within this setting, there is considerable success in optimizing secondary orthogonal transforms. The most prominent example is the Low-Frequency Non-Separable Transform (LFNST) in VVC. However, training nonlinear transforms without re-designing the RDOQ and entropy coding stage is a hard problem. In learned image compression, variational autoencoders have shown impressive results, but they use their own entropy model, remain difficult to train for small blocks and RDOQ is nontrivial for them. This paper describes a novel design of a nonlinear transform network for block-based video coding. Given a transform block, a fully-connected neural network predict coefficients from previously reconstructed ones and the adherent block boundary, such that only the residual coefficients need to be transmitted. Furthermore, another neural network filters the entire transform block before the inverse transform is applied and the intra prediction signal is added. Against the Versatile Video Coding Test Model 14.2 (VTM-14.2), luma bit-rate savings of approximately 1.9 % are reported for the All-Intra configuration.
Michael Schäfer 0003, Jonathan Pfaff, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
PCS1
2023 A Study on Data-Driven Probability Estimator Design for Video Coding
abstract
Data-driven optimization is employed to study alternative approaches [1] to the probability estimator of the the Enhanced Compression Model (ECM) (which includes additional coding tools on top of the Versatile Video Coding standard). In ECM, each context model uses a weighted sum of two hypotheses for probability estimation with different associated adaptation rates. Four alternative approaches are studied:
Heiner Kirchhoffer, Christian Rudat, Michael Schäfer 0003, Jonathan Pfaff, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
DCC3
2023 Block-Based Motion Estimation for Deep-Learned Video Coding
abstract
The research on deep-learned end-to-end video compression has attracted a lot of attention over the course of recent years. A central component of many approaches is to perform motion-compensated prediction by using convolutional neural networks (CNN) which determine a compressed representation of the motion field as features. Often, this task is divided into searching motion vectors by one network and efficiently representing them by another one. However, these networks may find motion fields far from optimal because the search radius of CNNs is mainly determined by their depth and kernel size. In this paper, we apply motion estimation techniques from classical block-based hybrid video compression to search a motion field which is then fed into a variational autoencoder. These strategies include different distortion measures, different block partitions and an improved approximation of the residual bitrate. With our modifications, bitrate savings of up to 13% over the underlying end-to-end based video codec can be obtained.
Sophie Pientka, Michael Schäfer 0003, Jonathan Pfaff, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
ICIP2
2022 Trellis-Coded Quantization for End-to-End Learned Image Compression
abstract
The performance of variational auto-encoders (VAE) for image compression has steadily grown in recent years, thus becoming competitive with advanced visual data compression technologies. These neural networks transform the source image into a latent space with a channel-wise representation. In most works, the latents are scalar quantized before being entropy coded. On the other hand, vector quantizers generally achieve denser packings of high-dimensional data regardless of the source distribution. Hence, low-complexity variants of these quantizers are implemented in the compression standards JPEG 2000 and Versatile Video Coding. In this paper we demonstrate coding gains by using trellis-coded quantization (TCQ) over scalar quantization. For the optimization of the networks with regard to TCQ, we employ a specific noisy representation of the features during the training stage. For variable-rate VAEs, we obtained 7.7% average BD-rate savings on the Kodak images by using TCQ over scalar quantization. When different networks per target bitrate are optimized, we report a relative coding gain of 2.4% due to TCQ.
Karsten Sühring, Michael Schäfer 0003, Jonathan Pfaff, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
ICIP2
2022 Deep video coding with gradient-descent optimized motion compensation and Lanczos filtering
abstract
Variational autoencoders have shown promising results for still image compression and have gained a lot of consideration in this field. Recently, noteworthy attempts were made to extend such end-to-end methods to the setting of video compression. Here, low-latency scenarios have been commonly investigated. In this paper, it is shown that the compression efficiency in this setting is improved by applying tools that are typically used in block-based hybrid coding such as rate-distortion optimized encoding of the features and advanced interpolation filters for computing samples at fractional positions. Additionally, a separate motion estimation network is trained to further increase the compression efficiency. Experimental results show that the rate-distortion performance benefits from including the aforementioned tools.
Sophie Pientka, Michael Schäfer 0003, Jonathan Pfaff, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
PCS2
2021 Rate-Distortion-Optimization for Deep Image Compression
abstract
Given the capabilities of massive GPU hardware, there has been a surge of using artificial neural networks (ANN) for still image compression. These compression systems usually consist of convolutional layers and can be considered as non-linear transform coding. Notably, these ANNs are based on an end-to-end approach where the encoder determines a compressed version of the image as features. In contrast to this, existing image and video codecs employ a block-based architecture with signal-dependent encoder optimizations. A basic requirement for designing such optimizations is estimating the impact of the quantization error on the resulting bitrate and distortion. As for non-linear, multi-layered neural networks, this is a difficult problem. This paper presents a performant auto-encoder architecture for still image compression, which represents the compressed features at multiple scales. Then, we demonstrate how an algorithm, which tests multiple feature candidates, can reduce the Lagrangian cost and optimize compression efficiency. The algorithm avoids multiple network executions by pre-estimating the impact of the quantization on the distortion by a higher-order polynomial.
Michael Schäfer 0003, Sophie Pientka, Jonathan Pfaff, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
ICIP1
2021 Extension of Matrix-Based Intra Prediction to 4: 4: 4 Chroma Formats
abstract
The final version of the Versatile Video Coding (VVC) standard incorporates the Matrix-Based Intra Prediction (MIP) tool. It consists of additional intra prediction modes which were derived from a data-driven training. These modes, in general, are applied to the luma component only. This paper describes how to apply the MIP modes to the chroma components in certain cases. In the generic case, if a chroma block uses the intra direct mode (DM), its intra mode is derived from the co-located luma block. If this luma block uses MIP, the chroma block uses the planar mode, because a MIP mode is not applicable for all block shapes. However, if the chroma format is 4:4:4 and single-tree coding is used, luma and chroma blocks share the same partitioning. The paper investigates the impact of using the MIP mode of the luma component when the chroma mode is the DM in these cases. This extension also harmonizes MIP with the Adaptive Color-Space Transform (ACT) for RGB-content, because the chroma mode of blocks using the ACT is inferred to be the DM. For the extended MIP tool in the Versatile Video Coding Test Model 11.2 (VTM-11.2) that implements the final VVC standard, bit-rate savings of 1.51%/0.60%/0.62% in terms of the Bj⊘ntegaard-Delta bit rate (BD-rate) are reported compared to the VTM-11.2 with MIP disabled for the All-Intra configuration and natural content in the RGB 4:4:4 format.
Björn Stallenberger, Michael Schäfer 0003, Philipp Merkle, Jonathan Pfaff, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
PCS2
2020 Efficient Fixed-Point Implementation Of Matrix-Based Intra Prediction
abstract
In the current version of the evolving Versatile Video Coding (VVC) standard, the Matrix-Based Intra Prediction (MIP) tool is included which is based on data-driven training techniques. This paper describes how an efficient fixed point implementation of the MIP prediction modes was designed. First, a method is presented that improves the integer quantization of the floating-point matrices resulting from the training by geometrically transforming the input and output. Second, a soft-clipping function can be used during the training stage to restrict the range of floating-point matrix entries. This enables the quantization of the trained coefficients with a fixed bit depth and precision. The clipping itself has no impact on the compression efficiency. The obtained predictors are incorporated into the Versatile Video Coding Test Model 7. All-Intra bit-rate savings of 0.6 % across different resolutions in terms of the Bjøntegaard-Delta bit rate (BD-rate) are reported.
Michael Schäfer 0003, Björn Stallenberger, Jonathan Pfaff, Philipp Helle, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
ICIP1
2020 A signal adaptive diffusion filter for video coding: Mathematical framework and complexity reductions
abstract
In this paper we combine video compression and modern image processing methods. We construct novel iterative filter methods for prediction signals based on Partial Differential Equation (PDE) based methods. The mathematical framework of the employed diffusion filter class is given and some desirable properties are stated. In particular, two types of diffusion filters are constructed: a uniform diffusion filter using a fixed filter mask and a signal adaptive diffusion filter that incorporates the structures of the underlying prediction signal. The latter has the advantage of not attenuating existing edges while the uniform filter is less complex. The filters are embedded into a software based on HEVC with additional QTBT (Quadtree plus Binary Tree) and MTT (Multi-Type-Tree) block structure. In this setting, several measures to reduce the coding complexity of the tool are introduced, discussed and tested thoroughly. The coding complexity is reduced by up to 70% while maintaining over 80% of the gain. Overall, the diffusion filter method achieves average bitrate savings of 2.27% for Random Access having an average encoder runtime complexity of 119% and 117% decoder runtime complexity. For individual test sequences, results of 7.36% for Random Access are accomplished.
Jennifer Rasch, Jonathan Pfaff, Michael Schäfer 0003, Anastasia Henkel, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
Signal Process. Image Commun.3
2020 Video Compression Using Generalized Binary Partitioning, Trellis Coded Quantization, Perceptually Optimized Encoding, and Advanced Prediction and Transform Coding
abstract
In this paper, we describe a video coding design that enables a higher coding efficiency than the HEVC standard. The proposed video codec follows the design of block-based hybrid video coding, but includes a number of advanced coding tools. A part of the incorporated advanced concepts was developed by the Joint Video Exploration Team, while others are newly proposed. The key aspects of these newly proposed tools are the following. A video frame is subdivided into rectangles of variable size using a binary partitioning with variable split ratios. Three new approaches for generating spatial intra prediction signals are supported: A line-wise application of conventional intra prediction modes, coupled with a mode-dependent processing order, a region-based template matching prediction method and intra prediction modes based on neural networks. For motion-compensated prediction, a multi-hypothesis mode with more than two motion hypotheses can be used. In transform coding, mode dependent combinations of primary and secondary transforms are applied. Moreover, scalar quantization is replaced by trellis-coded quantization and the entropy coding of the quantized transform coefficients is improved. The intra and inter prediction signals can be filtered using an edge-preserving diffusion filter or a non-linear DCT-based thresholding operation. The video codec includes an adaptive in-loop filter for which one of three classifiers can be chosen on a picture basis. We also incorporated an optional encoder control, which adjusts the quantization parameters based on a perceptually motivated distortion measure. In a random access scenario, our proposed video codec achieves luma BD-rate savings between 32.5% for HDR HLG UHD and 39.6% for SDR UHD over the HEVC (HM software) anchor for different categories of test sequences.
Jonathan Pfaff, Heiko Schwarz, Detlev Marpe, Benjamin Bross, Santiago De-Luxán-Hernández, Philipp Helle, Christian R. Helmrich, Tobias Hinz, Wang-Q Lim, Jackie Ma, Tung Nguyen 0001, Jennifer Rasch, Michael Schäfer 0003, Mischa Siekmann, Gayathri Venugopal, Adam Wieckowski, Martin Winken, Thomas Wiegand 0001
IEEE Trans. Circuits Syst. Video Technol.13
2019 Intra Picture Prediction for Video Coding with Neural Networks
abstract
We train a neural network to perform intra picture prediction for block based video coding. Our network has multiple prediction modes which co-adapt during training to minimize a loss function. By applying the l1-norm and a sigmoid-function to the prediction residual in the DCT domain, our loss function reflects properties of the residual quantization and coding stages present in the typical hybrid video coding architecture. We simplify the resulting predictors by pruning them in the frequency domain, thus greatly reducing the number of multiplications otherwise needed for the dense matrix-vector multiplications. Also, by quantizing the network weights and using fixed point arithmetic, we allow for a hardware friendly implementation. We demonstrate significant coding gains over state of the art intra prediction.
Philipp Helle, Jonathan Pfaff, Michael Schäfer 0003, Roman Rischke, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
DCC3
2019 An Affine-Linear Intra Prediction With Complexity Constraints
abstract
This paper presents a novel method for a data-driven training of affine-linear predictors which perform intra prediction in state-of-the-art video coding. The main aspect of our training design is the use of subband decomposition of both the input and the output of the prediction. Due to this architecture, the same set of predictors can be shared across different block shapes leading to a very limited memory requirement. Also, the computational complexity of the resulting predictors can be limited such that it does not exceed the complexity of the conventional angular intra prediction. In the training itself, a loss function modelling the bit-rate of the DCT-transformed residuals is used. The obtained predictors are incorporated into the Versatile Video Coding Test Model 3 in addition to the conventional intra prediction modes. All-Intra bit-rate savings ranging from 0.8% to 1.4% across different resolutions have been measured in terms of the Bjøntegaard-Delta bit rate (BD-rate).
Michael Schäfer 0003, Björn Stallenberger, Jonathan Pfaff, Philipp Helle, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
ICIP1
2019 A Data-Trained, Affine-Linear Intra-Picture Prediction in the Frequency Domain
abstract
This paper presents a data-driven training of affine- linear predictors which perform intra-picture prediction for video coding. The trained predictors use a single line of reconstructed boundary samples as input like the conventional intra prediction modes. For large blocks, the presented predictors initially transform the input samples via Discrete Cosine Transform. This allows to omit high frequency coefficients and consequently reduce the input dimension. The output is the result of a single matrix-vector multiplication and offset addition. Here, the predictors only construct certain coefficients in the frequency domain. The final prediction signal is then obtained by inverse transform. The coefficients of the prediction modes need to be stored in advance, requiring 0.273 MB of memory. The training employs a recursive block partitioning, where the loss function targets to approximate the bit-rate of the DCT-transformed block residuals. The obtained predictors are incorporated into the Versatile Video Coding Test Model 4. The authors report All- Intra bit-rate savings ranging from 0.7% to 2.0% across different resolutions in terms of the Bjøntegaard-Delta bit rate (BD-rate).
Michael Schäfer 0003, Björn Stallenberger, Jonathan Pfaff, Philipp Helle, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001
PCS1
2018 Improved Prediction Via Thresholding Transform Coefficients
abstract
This paper presents a thresholding method for processing the predicted samples in the state-of-the-art High Efficiency Video Coding (HEVC) standard. The method applies an integer-based approximation of the discrete cosine transform to an extended prediction block and sets transform coefficients beneath a certain threshold to zero. Transforming back into the sample domain yields the improved prediction signal. The method is incorporated into a software implementation that is conforming to the HEVC standard and applies to both intra and inter predictions. Consequently, bit-rate savings ranging from 2.3% to 8.0% have been measured in terms of the Bjøntegaard-Delta bit rate (BD-rate).
Michael Schäfer 0003, Jonathan Pfaff, Jennifer Rasch, Tobias Hinz, Heiko Schwarz, Tung Nguyen 0001, Gerhard Tech, Detlev Marpe, Thomas Wiegand 0001
ICIP1
2018 A Signal Adaptive Diffusion Filter For Video Coding
abstract
In this paper we combine state of the art video compression and Partial Differential Equation (PDE) based image processing methods. We introduce a new signal adaptive method to filter the predictions of a hybrid video codec using a system of PDEs describing a diffusion process. The method can be applied to intra as well as inter predictions. The filter is embedded into the framework of HEVC. The efficiency of the HEVC video codec is improved by up to -2.76% for All Intra and -3.56% for Random Access measured in Bjøntegaard delta (BD) rate. Coding gains of up to -8.76% can be observed for individual test sequences.
Jennifer Rasch, Jonathan Pfaff, Michael Schäfer 0003, Heiko Schwarz, Martin Winken, Mischa Siekmann, Detlev Marpe, Thomas Wiegand 0001
PCS3