VLDB 2026 Research / reviewers in the wild / expert
Gerhard Tech
dblp:10/8844
· DBLP profile ↗
15ranked-venue papers
9as first author
5since 2021 · last 2024
0000-0003-0072-2284ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 9 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Neural Network Coding of Difference Updates for Efficient Distributed Learning CommunicationabstractDistributed learning requires a frequent communication of neural network update data. For this, we present a set of new compression tools, jointly called differential neural network coding (dNNC). dNNC is specifically tailored to efficiently code incremental neural network updates and includes tools for federated BatchNorm folding (FedBNF), structured and unstructured sparsification, tensor row skipping, quantization optimization and temporal adaptation for improved context-adaptive binary arithmetic coding (CABAC). Furthermore, dNNC provides a new parameter update tree (PUT) mechanism, which allows to identify updates for different neural network parameter sub-sets and their relationship in synchronous and asynchronous neural network communication scenarios. Most of these tools have been included into the standardization process of the NNC standard (ISO/IEC 15938-17) edition 2. We benchmark dNNC in multiple federated and split learning scenarios using a variety of NN models and data including vision transformers and large-scale ImageNet experiments: It achieves compression efficiencies of 60% in comparison to the NNC standard edition 1 for transparent coding cases, i.e., without degrading the inference or training performance. This corresponds to a reduction in the size of the NN updates to less than 1% of their original size. Moreover, dNNC reduces the overall energy consumption required for communication in federated learning systems by up to 94%. Daniel Becking, Karsten Müller 0001, Paul Haase, Heiner Kirchhoffer, Gerhard Tech, Wojciech Samek, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | History Dependent Significance Coding for Incremental Neural Network CompressionabstractThis paper presents an improved probability estimation scheme for the entropy coder of Incremental Neural Network Coding (INNC), which is currently under standardization in ISO/IEC MPEG. More specifically, the paper first analyzes the compression performance of INNC and how the bitstream size relates to the neural network (NN) layers. For the layers requiring the most bits, it analyzes the coded NN weight updates and their temporal dependencies. Major finding is that the probability of a significant (i.e., non-zero) update for a weight can depend considerably on whether the weight has been updated before. Based on this finding, the paper proposes a new probability estimation scheme: Depending on whether a significant update has been received before (i.e., based on the weight’s history), the entropy coder models the probability for a current significant update differently. This scheme achieves a bitstream size reduction of about 2% and 1% in a transfer and a federated learning scenario, respectively, without any accuracy loss or significant complexity increase. Therefore, MPEG adopted our history dependent significance probability (HDSP) scheme to its emerging standard for INNC. Gerhard Tech, Paul Haase, Daniel Becking, Heiner Kirchhoffer, Karsten Müller 0001, Jonathan Pfaff, Heiko Schwarz, Wojciech Samek, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 1 |
| 2021 | Fast Partitioning for VVC Intra-Picture Encoding with a CNN Minimizing the Rate-Distortion-Time CostabstractThis paper presents a CNN to reduce the encoding time of a VVC-based intra-picture encoder. For encoding a 32 × 32 block, the CNN estimates two partitioning parameters that restrict the allowed coding block width and height. To estimate them such that the encoder skips testing inefficient partitioning modes, we train the CNN as follows: First, we generate training data by encoding sequences without the CNN. While encoding, we test all combinations of the two parameters for each 32 × 32 block and store the resulting Lagrangian rate-distortion-time (RDT) cost. We use the recorded cost to derive the loss function when training the CNN. Consequently, the CNN is trained such that it minimizes the Lagrangian RDT cost. Our CNN reduces the encoding time by 50% with a bit rate increase of 0.9%, which outperforms existing CNN-based approaches. Our generic training approach could also be applied for other encoder parameters. Gerhard Tech, Jonathan Pfaff, Heiko Schwarz, Philipp Helle, Adam Wieckowski, Detlev Marpe, Thomas Wiegand 0001 |
DCC | 1 |
| 2021 | CNN-based parameter selection for fast VVC intra-picture encodingabstractThis paper presents two new methods for fast VVC intra-picture encoding. Both are based on an approach that uses a CNN for blockadaptive parameter estimation. The parameters restrict the multitype-tree (MTT) partitionings tested by the encoder. The methods aim for an improvement of the approach by further constraints with additional parameters. Adding parameters increases the time required for training data generation exponentially. This raises the question which parameters to add and how. To explore further partitioning restrictions, the first method adds parameters controlling the block sizes the MTT can start from. Although this leads to four parameters, we can exploit that some of their combinations are invalid. To investigate whether testing fewer prediction and transform modes is feasible, the second method adds a single parameter that restricts their number jointly. The paper evaluates hypothetical and actual encoding time reductions for VTM-10.2. The first method outperforms our other and other existing method: The encoding time decreases by 50% with a bit rate increase of 0.7%. Gerhard Tech, Jonathan Pfaff, Heiko Schwarz, Philipp Helle, Adam Wieckowski, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 1 |
| 2021 | Rate-Distortion-Time Cost Aware CNN Training for Fast VVC Intra-Picture Partitioning DecisionsabstractThis paper presents a new method for fast VVC intra-picture encoding using a CNN. The CNN operates on the original samples of$\mathbf{32}\times \mathbf{32}$blocks. Given a current block, it derives for each of the block's multi-type trees (MTTs), which are nested in quad-tree (QT) nodes, a parameter pair. The parameter pairs constrain the minimum width and height of the sub-blocks in their MTTs. This enables the CNN to control the number of tested MTT splits with fine granularity. To skip modes while maintaining the rate-distortion (RD) performance, we train the CNN considering the Lagrangian rate-distortion-time (RDT) cost caused by the derived parameters. First, we generate training data by encoding; when reaching a quad-tree node in a$\mathbf{32}\times \mathbf{32}$block, we encode the associated MTT with varying parameter pair values and record the resulting the RD and time cost. Then, when the CNN outputs parameters in training, we estimate the related RDT cost of the$\mathbf{32}\times \mathbf{32}$block using the recorded data. For this, we model the dependency between RDT cost and the parameters by emulating the encoder's RD optimization process. This way, we train the CNN while considering the RDT cost with an accuracy that is sufficient to outperform existing approaches. The approach achieves an encoding time reduction of 50% with a bit rate increase of only 0.7% for VTM-10.2. Gerhard Tech, Jonathan Pfaff, Heiko Schwarz, Philipp Helle, Adam Wieckowski, Detlev Marpe, Thomas Wiegand 0001 |
PCS | 1 |
| 2018 | Improved Prediction Via Thresholding Transform CoefficientsabstractThis paper presents a thresholding method for processing the predicted samples in the state-of-the-art High Efficiency Video Coding (HEVC) standard. The method applies an integer-based approximation of the discrete cosine transform to an extended prediction block and sets transform coefficients beneath a certain threshold to zero. Transforming back into the sample domain yields the improved prediction signal. The method is incorporated into a software implementation that is conforming to the HEVC standard and applies to both intra and inter predictions. Consequently, bit-rate savings ranging from 2.3% to 8.0% have been measured in terms of the Bjøntegaard-Delta bit rate (BD-rate). Michael Schäfer 0003, Jonathan Pfaff, Jennifer Rasch, Tobias Hinz, Heiko Schwarz, Tung Nguyen 0001, Gerhard Tech, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 7 |
| 2018 | Partial Depth Image Based Re-Rendering for Synthesized View Distortion Computationabstract3D video systems transmit depth maps in order to render synthesized views (SVs) at a receiver. To anticipate this purpose when processing a depth map, a sender-side depth processing algorithm (DPA), e.g. a depth encoder, can also render the SVs, compute their SV distortion (SVD), and adapt to it. This requires a low-complexity algorithm as computational resources are usually limited. We propose such an algorithm in this paper. First, we discuss a measure that relates a depth change to an SVD change using rendering. Then, we present an optimized process combining basic rendering steps, as warping, occlusion handling, interpolation, hole filling, and blending. Furthermore, we analyze which parts of an SV are affected by a depth change and modify the process to re-render only them. The resulting algorithm is significantly less complex than an unoptimized rendering-based variant and quantifies the SVD more accurately than existing estimation methods. The algorithm is used by the 3D-High Efficiency Video Coding reference software encoder as the main method for distortion computation and can also be used by other DPAs. Gerhard Tech, Karsten Müller 0001, Heiko Schwarz, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Overview of the Multiview and 3D Extensions of High Efficiency Video CodingabstractThe High Efficiency Video Coding (HEVC) standard has recently been extended to support efficient representation of multiview video and depth-based 3D video formats. The multiview extension, MV-HEVC, allows efficient coding of multiple camera views and associated auxiliary pictures, and can be implemented by reusing single-layer decoders without changing the block-level processing modules since block-level syntax and decoding processes remain unchanged. Bit rate savings compared with HEVC simulcast are achieved by enabling the use of inter-view references in motion-compensated prediction. The more advanced 3D video extension, 3D-HEVC, targets a coded representation consisting of multiple views and associated depth maps, as required for generating additional intermediate views in advanced 3D displays. Additional bit rate reduction compared with MV-HEVC is achieved by specifying new block-level video coding tools, which explicitly exploit statistical dependencies between video texture and depth and specifically adapt to the properties of depth maps. The technical concepts and features of both extensions are presented in this paper. Gerhard Tech, Ying Chen 0011, Karsten Müller 0001, Jens-Rainer Ohm, Anthony Vetro, Ye-Kui Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | 3D High-Efficiency Video Coding for Multi-View Video and Depth DataabstractThis paper describes an extension of the high efficiency video coding (HEVC) standard for coding of multi-view video and depth data. In addition to the known concept of disparity-compensated prediction, inter-view motion parameter, and inter-view residual prediction for coding of the dependent video views are developed and integrated. Furthermore, for depth coding, new intra coding modes, a modified motion compensation and motion vector coding as well as the concept of motion parameter inheritance are part of the HEVC extension. A novel encoder control uses view synthesis optimization, which guarantees that high quality intermediate views can be generated based on the decoded data. The bitstream format supports the extraction of partial bitstreams, so that conventional 2D video, stereo video, and the full multi-view video plus depth format can be decoded from a single bitstream. Objective and subjective results are presented, demonstrating that the proposed approach provides 50% bit rate savings in comparison with HEVC simulcast and 20% in comparison with a straightforward multi-view extension of HEVC without the newly developed coding tools. Karsten Müller 0001, Heiko Schwarz, Detlev Marpe, Christian Bartnik, Sebastian Bosse, Heribert Brust, Tobias Hinz, Haricharan Lakshman, Philipp Merkle, Hunn Rhee, Gerhard Tech, Martin Winken, Thomas Wiegand 0001 |
IEEE Trans. Image Process. | 11 |
| 2012 | Extension of High Efficiency Video Coding (HEVC) for multiview video and depth dataabstractThis paper presents an approach for 3D video coding that uses a format in which a small number of views as well as associated depth maps are coded and transmitted. At the receiver side, additional views required for displaying the 3D video on an autostereoscopic display can be generated based on the corresponding decoded signals by using depth image based rendering (DIBR) techniques. In terms of coding technology, the proposed coding scheme represents an extension of High Efficiency Video Coding (HEVC), similar to the Multiview Coding (MVC) extension of H.264/AVC. Besides the well-known disparity-compensated prediction, advanced techniques for inter-view and inter-component prediction, the representation of depth blocks, and the encoder control for depth signals have been developed and integrated. In comparison to simulcasting the different signals using HEVC, the proposed approach provides about 40% and 50% average bit rate savings for a whole test set when configured to comply with a 2- and 3-view scenario, respectively. The proposed codec was submitted as response to a Call for Proposals on 3D Video Technology issued by the ISO/IEC Moving Picture Experts Group (MPEG) and it was ranked as the overall best performing HEVC-based proposal in the related subjective tests. Heiko Schwarz, Christian Bartnik, Sebastian Bosse, Heribert Brust, Tobias Hinz, Haricharan Lakshman, Philipp Merkle, Karsten Müller 0001, Hunn Rhee, Gerhard Tech, Martin Winken, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 10 |
| 2012 | Synthesized View Distortion Based 3D Video Coding for Extrapolation and Interpolation of ViewsabstractIn 3D video coding, the Multi-View Video plus Depth (MVD) representation format enables the interpolation of virtual intermediate views in between the original recorded camera views, as well as the extrapolation of views outside this range using view synthesis algorithms. For this, MVD includes depth maps as additional information to be transmitted. The performance of the associated depth data coding can be improved by considering the distortion of synthesized views in the rate-distortion optimization of the encoder. Therefore, this paper evaluates coding gains achieved with encoder optimization with respect to virtual view positions. The impact of the number and the positions of virtual views used as reference in the encoder side optimization process on the quality of the synthesized views is analyzed. Coding results for view interpolation and view extrapolation are provided. Moreover, the encoder-side complexity with respect to the number of reference views is discussed. Gerhard Tech, Heiko Schwarz, Karsten Müller 0001, Thomas Wiegand 0001 |
ICME | 1 |
| 2012 | 3D video coding using advanced prediction, depth modeling, and encoder control methodsabstractThe presented approach for 3D video coding uses the multiview video plus depth format, in which a small number of video views as well as associated depth maps are coded. Based on the coded signals, additional views required for displaying the 3D video on an autostereoscopic display can be generated by depth image based rendering techniques. The developed coding scheme represents an extension of HEVC, similar to the MVC extension of H.264/AVC. However, in addition to the well-known disparity-compensated prediction advanced techniques for inter-view and inter-component prediction, the representation of depth blocks, and the encoder control for depth signals have been integrated. In comparison to simulcasting the different signals using HEVC, the proposed approach provides about 40% and 50% bit rate savings for the tested configurations with 2 and 3 views, respectively. Bit rate reductions of about 20% have been obtained in comparison to a straightforward multiview extension of HEVC without the newly developed coding tools. Heiko Schwarz, Christian Bartnik, Sebastian Bosse, Heribert Brust, Tobias Hinz, Haricharan Lakshman, Detlev Marpe, Philipp Merkle, Karsten Müller 0001, Hunn Rhee, Gerhard Tech, Martin Winken, Thomas Wiegand 0001 |
PCS | 11 |
| 2012 | 3D video coding using the synthesized view distortion changeabstractIn 3D video, texture and supplementary depth data are coded to enable the interpolation of a required number of synthesized views for multi-view displays in the range of the original camera views. The coding of the depth data can be improved by analyzing the distortion of synthesized video views instead of the depth map distortion. Therefore, this paper introduces a new distortion metric for 3D video coding, which relates changes in the depth map directly to changes of the overall synthesized view distortion. It is shown how the new metric can be integrated into the rate-distortion optimization (RDO) process of an encoder, that is based on high-efficiency video coding technology. An evaluation of the modified encoder is conducted using different view synthesis algorithms and shows about 50% rate savings for the depth data or 0.6 dB PSNR gains for the synthesized view. Gerhard Tech, Heiko Schwarz, Karsten Müller 0001, Thomas Wiegand 0001 |
PCS | 1 |
| 2010 | 3D video formats and coding methodsabstractThe introduction of first 3D systems for digital cinema and home entertainment is based on stereo technology. For efficiently supporting new display types, depth-enhanced formats and coding technology is required, as introduced in this overview paper. First, we discuss the necessity for a generic 3D video format, as the current state-of-the-art in multi-view video coding cannot support different types of multi-view displays at the same time. Therefore, a generic depth-enhanced 3D format is developed, where any number of views can be generated from one bit stream. This, however, requires a complex framework for 3D video, where not only the 3D format and new coding methods are investigated, but also view synthesis and the provision of high-quality depth maps, e.g. via depth estimation. We present this framework and discuss the interdependencies between the different modules. Karsten Müller 0001, Philipp Merkle, Gerhard Tech, Thomas Wiegand 0001 |
ICIP | 3 |
| 2010 | Diffusion filtering of depth maps in stereo video codingabstractA method for removing irrelevant information from depth maps in Video plus Depth coding is presented. The depth map is filtered in several iterations using a diffusional approach. In each iteration smoothing is carried out in local sample neighborhoods considering the distortion introduced to a rendered view. Smoothing is only applied when the rendered view is not affected. Therefore irrelevant edges and features in the depth map can be damped while the quality of the rendered view is retained. The processed depth maps can be coded at a reduced rate compared to unaltered data. Coding experiments show gains up to 0.5 dB for the rendered view at the same bit rate. Gerhard Tech, Karsten Müller 0001, Thomas Wiegand 0001 |
PCS | 1 |