EDBT 2026 Demo / reviewers in the wild / expert
Tobias Hinz
dblp:08/645
· DBLP profile ↗
38ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-level Inter-frame Parallelization in an Open Optimized VVC EncoderabstractThis work investigates video encoding parallelization techniques based on the Versatile Video Coding (VVC) standard, using the open and optimized encoder software implementation VVenC. Modern multi-processor systems offer significant opportunities for accelerating video encoding. By employing a proposed combination of parallelization methods, the VVenC encoder achieves an acceleration factor of up to 22 compared to single-threaded mode on a 32-core system, with potential increases to 27× at higher bitrates. Building upon prior work on Inter-frame Parallelization (IFP), the study introduces frame region-based synchronization, enabling further acceleration of up to 10%. Beyond that, the study demonstrates extending frame parallelization beyond Group of Pictures (GOP) boundaries, which improves IFP speed up by 37% and 11% at high-definition (HD) and ultra-high-definition (UHD) resolutions, respectively. Additional combinations with other VVC parallelization tools, such as tiles and VVC Wavefront Parallel Processing (WPP), are also explored. The article provides a comprehensive analysis of parallelization challenges and highlights areas for further improvement. Valeri George, Jens Brandenburg, Gabriel Hege, Tobias Hinz, Adam Wieckowski, Benjamin Bross, Thomas Schierl, Detlev Marpe |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion ModelsabstractCurrent diffusion-based text-to-video methods are limited to producing short video clips of a single shot and lack the capability to generate multi-shot videos with discrete transitions where the same character performs distinct activities across the same or different backgrounds. To address this limitation we propose a framework that includes a dataset collection pipeline and architectural extensions to video diffusion models to enable text-to-multi-shot video generation. Our approach enables generation of multi-shot videos as a single video with full attention across all frames of all shots, ensuring character and background consistency, and allows users to control the number, duration, and content of shots through shot-specific conditioning. This is achieved by incorporating a transition token into the text-to-video model to control at which frames a new shot begins and a local attention masking strategy which controls the transition token’s effect and allows shot-specific prompting. To obtain training data we propose a novel data collection pipeline to construct a multi-shot video dataset from existing single-shot video datasets. Extensive experiments demonstrate that fine-tuning a pre-trained text-to-video model for a few thousand iterations is enough for the model to subsequently be able to generate multi-shot videos with shot-specific control, outperforming the baselines. You can find more details in our webpage. Özgür Kara, Krishna Kumar Singh, Duygu Ceylan, James M. Rehg, Tobias Hinz |
CVPR | 6 |
| 2025 | Digital signatures for trustworthy authentication of elementary video streamsabstractThis paper introduces a method for digitally signing and verifying elementary video bitstreams. The method verifies temporal consistency of the video while allowing random access into the bitstream and adaptation to temporal and spatial scalability including sub-bitstream extraction. It was adopted by the Joint Video Experts Team (JVET) into the Versatile supplemental enhancement information messages for coded video bitstreams (VSEI) specification version 4 by introducing three new Digitally Signed Content SEI messages. Karsten Sühring, Tobias Hinz, Jonathan Pfaff, Yago Sánchez de la Fuente, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
VCIP | 2 |
| 2024 | Personalized Residuals for Concept-Driven Text-to-Image GenerationabstractWe present personalized residuals and localized attention-guided sampling for efficient concept-driven generation using text-to-image diffusion models. Our method first represents concepts by freezing the weights of a pretrained text-conditioned diffusion model and learning low-rank residuals for a small subset of the model's layers. The residual-based approach then directly enables application of our proposed sampling technique, which applies the learned residuals only in areas where the concept is localized via cross-attention and applies the original diffusion weights in all other regions. Localized sampling therefore combines the learned identity of the concept with the existing generative prior of the underlying diffusion model. We show that personalized residuals effectively capture the identity of a concept in$\sim$3 minutes on a single GPU without the use of regularization images and with fewer parameters than previous models, and localized sampling allows using the original model as strong prior for large parts of the image. Cusuh Ham, Matthew Fisher, James Hays, Nicholas I. Kolkin, Yuchen Liu 0002, Richard Zhang 0001, Tobias Hinz |
CVPR | 7 |
| 2024 | SNED: Superposition Network Architecture Search for Efficient Video Diffusion ModelabstractWhile AI-generated content has garnered significant attention, achieving photo-realistic video synthesis remains a formidable challenge. Despite the promising advances in diffusion models for video generation quality, the complex model architecture and substantial computational demands for both training and inference create a significant gap between these models and real-world applications. This paper presents SNED, a superposition network architecture search method for efficient video diffusion model. Our method employs a supernet training paradigm that targets various model cost and resolution options using a weight-sharing method. Moreover, we propose the supernet training sampling warm-up for fast training optimization. To showcase the flexibility of our method, we conduct experiments involving both pixel-space and latent-space video diffusion models. The results demonstrate that our framework consistently produces comparable results across different model options with high efficiency. According to the experiment for the pixel-space video diffusion model, we can achieve consistent video generation results simultaneously across 64×64 to 256×256 resolutions with a large range of model sizes from 640M to 1.6B number of parameters for pixel-space video diffusion models. Zhengang Li 0001, Yuchen Liu 0002, Difan Liu, Tobias Hinz, Feng Liu 0015, Yanzhi Wang 0001 |
CVPR | 5 |
| 2024 | Inter-Frame Parallelization in an Open Optimized VVC EncoderabstractThe Versatile Video Coding (VVC) standard promises high compression efficiency for diverse content types. Based on VVenC, an open and optimized VVC software video encoder, this work presents an inter-frame parallelization (IFP) method designed to exploit the processing power of modern platforms featuring a high number of computing cores. Encoding an ultrahigh definition video on a 32-core machine with the VVenC's faster preset, the proposed method shows more than 20% increase in encoder speed while only a 1% decrease in compression efficiency compared to the default multi-threading mode. In comparison to single-threaded mode, it corresponds to a speedup factor of 18, up from 15x achievable with the previous parallelization scheme. Furthermore, the synergy of the developed inter-frame parallelization technique with other parallelization methods is explored, including tiles and VVC wavefront parallel processing (WPP). The combination of these approaches enables a notable speedup factor of 21, albeit with a trade-off in coding efficiency. With a focus on VVC, this research contributes to the ongoing discourse on video coding optimization, providing valuable insights into possible pitfalls and the potential gains achievable through efficient parallelization techniques on high-core platforms. Valeri George, Jens Brandenburg, Gabriel Hege, Tobias Hinz, Adam Wieckowski, Benjamin Bross, Thomas Schierl, Detlev Marpe |
MMSys | 4 |
| 2024 | Fast First Pass in Two-Pass Video Encoding Using Sub-SamplingabstractRate control (RC), specifically two-pass, is the main operation mode in VVenC, an open and optimized Versatile Video Coding (VVC) encoder. VVC offers substantial bitrate savings over its predecessor, High Efficiency Video Coding (HEVC), at the price of increased complexity. This complexity increase is apparent in both encoding passes of VVenC. While the complexity redaction in the final pass has been discussed, this paper considers complexity reduction in the first pass, in addition to its already reduced search space. To reduce the overall runtime of a two-pass RC method, spatial and temporal sub-sampling of the first encoding pass is proposed. The experimental results show that the proposed first-pass sub-sampling in two-pass RC can speed up the encoding process of the default two-pass rate control algorithm in VVenC by 18%, with 0.48% loss in coding efficiency, when using the faster preset. Using temporal sub-sampling for the look-ahead, one-pass RC in VVenC can achieve time savings of 11% for bit-rate increases of 0.28%. Anastasia Henkel, Christian R. Helmrich, Tobias Hinz, Jens Brandenburg, Adam Wieckowski, Benjamin Bross, Detlev Marpe, Thomas Wiegand 0001 |
PCS | 3 |
| 2023 | SmartBrush: Text and Shape Guided Object Inpainting with Diffusion ModelabstractGeneric image inpainting aims to complete a corrupted image by borrowing surrounding information, which barely generates novel content. By contrast, multi-modal inpainting provides more flexible and useful controls on the inpainted content, e.g., a text prompt can be used to describe an object with richer attributes, and a mask can be used to constrain the shape of the inpainted object rather than being only considered as a missing area. We propose a new diffusion-based model named SmartBrush for completing a missing region with an object using both text and shape-guidance. While previous work such as DALLE-2 and Stable Diffusion can do text-guided inapinting they do not support shape guidance and tend to modify background texture surrounding the generated object. Our model incorporates both text and shape guidance with precision control. To preserve the background better, we propose a novel training and sampling strategy by augmenting the diffusion U-net with object-mask prediction. Lastly, we introduce a multi-task training strategy by jointly training inpainting with text-to-image generation to leverage more training data. We conduct extensive experiments showing that our model outperforms all baselines in terms of visual quality, mask controllability, and background preservation. Shaoan Xie, Zhe Lin 0001, Tobias Hinz, Kun Zhang 0001 |
CVPR | 4 |
| 2023 | Finalization of VVenC's Screen Content Detector and Two-Pass Rate Control Using Pre-Filtering StatisticsabstractFor improved performance, practical video encoders integrate algorithms for screen content detection and rate control. This paper outlines recently implemented optimizations to both the screen content classifier (SCC) and two-pass rate control (RC) of VVenC, an open Versatile Video Coding (VVC) compliant encoder. The improvements, confirmed by evaluation experiments in random-access configurations using an extended set of test videos, are mainly achieved by leveraging motion error statistics acquired during motion compensated temporal pre-filtering (MCTPF), carried out in VVenC’s pre-analysis stage. All three aspects – pre-analysis stage, SCC, and RC – are revisited herein, and the exploitation of MCTPF data is described. Christian R. Helmrich, Anastasia Henkel, Tobias Hinz, Adam Wieckowski, Benjamin Bross, Detlev Marpe |
ICIP | 3 |
| 2022 | Efficient Multi-Threading Strategies in VVenC, an Open and Optimized VVC Encoder ImplementationabstractThe Versatile Video Coding (VVC) standard has been developed to meet the ever-increasing demand for higher compression of digital video data. Compared to its predecessor, the High-Efficiency Video Coding (HEVC) standard, VVC reduces the bitrate by around 50% for the same perceived quality. This increase in compression efficiency is associated with an increase in computational complexity, mainly on the encoder side. As an open and optimized VVC software encoder implementation, VVenC integrates algorithmic optimizations for each coding tool in VVC. This allows to define a set of five presets from faster to slower as Pareto-optimal tradeoffs between runtime and efficiency. On top, multithreading allows to reduce the runtime and preserves most of the compression efficiency of each preset. This paper presents and analyses the different multi-threading strategies in VVenC. Using a combination of pre-processing, picture-level and in-picture parallelization, VVenC can achieve a parallelization speedup with a factor of 4 for 4 threads while reducing the compression efficiency by only 0.4%. For higher thread numbers, i.e. 16, the speedup depends on the video resolution and used encoder preset, ranging from 6-9 for high definition to 10-12 for ultrahigh definition video with similar loss of compression efficiency. Using additional wavefront and tiles in-picture parallelization, higher speedups can be achieved at the costs of decreased coding efficiency. Valeri George, Jens Brandenburg, Gabriel Hege, Tobias Hinz, Adam Wieckowski, Benjamin Bross, Detlev Marpe |
ISM | 4 |
| 2022 | A Scene Change and Noise Aware Rate Control Method for VVenC, An Open VVC Encoder ImplementationabstractContemporary motion picture content, consisting of scenes with different amounts of visual complexity or camera noise, represents demanding input for video encoders operating in rate control (RC) modes. This paper presents improvements to the 2-pass RC method integrated into VVenC, an open VVC encoder implementation, outlined in previous publications. We specifically introduce three extensions to our RC solution: first, frame type adaptation operating near scene cuts, along with an associated simple detector; second, rate stabilization means to allow for more reliable lookahead based 2-pass RC operation in on-the-fly encoding applications; and third, a low-complexity approach for estimating the instantaneous intensity of camera noise or film grain to avoid large variations in bit consumption when encoding individual frames in the final RC pass. Experimental evaluation confirms that these extensions significantly improve both the objective (BD rate) and subjective (visual) RC performance of VVenC especially on challenging video content. Christian R. Helmrich, Christian Bartnik, Jens Brandenburg, Valeri George, Tobias Hinz, Christian Lehmann, Ivan Zupancic, Adam Wieckowski, Benjamin Bross, Detlev Marpe |
PCS | 5 |
| 2022 | An Optimized Temporal Filter Implementation for Practical ApplicationsabstractVVenC, an open and optimized VVC encoder implementation, employs a temporal filter from the literature as a pre-processing step. The filter effectively reduces camera noise from input video, thereby increasing the encoding gain for lossy encoding, at a price of fairly high complexity, further increased by the necessity of consistent application to many pictures. The filter represents one of the most runtime consuming processing steps for the fastest operating points of VVenC. In this paper, steps are described to reduce the complexity of the temporal filtering in VVenC, to allow its application with low-complexity presets. Overall, the filter runtime is reduced by a factor of around 17 compared to the state of the art, while slightly improving its performance. An additional 4 times speedup is achieved using vectorized implementation. In the proposed version, for the VVenC preset faster, the filter provides 7.36% BD-rate gain at only 2% runtime overhead. Adam Wieckowski, Tobias Hinz, Christian R. Helmrich, Benjamin Bross, Detlev Marpe |
PCS | 2 |
| 2022 | CharacterGAN: Few-Shot Keypoint Character Animation and ReposingabstractWe introduce CharacterGAN, a generative model that can be trained on only a few samples (8 – 15) of a given character. Our model generates novel poses based on keypoint locations, which can be modified in real time while providing interactive feedback, allowing for intuitive reposing and animation. Since we only have very limited training samples, one of the key challenges lies in how to address (dis)occlusions, e.g. when a hand moves behind or in front of a body. To address this, we introduce a novel layering approach which explicitly splits the input keypoints into different layers which are processed independently. These layers represent different parts of the character and provide a strong implicit bias that helps to obtain realistic results even with strong (dis)occlusions. To combine the features of individual layers we use an adaptive scaling approach conditioned on all keypoints. Finally, we introduce a mask connectivity constraint to reduce distortion artifacts that occur with extreme out-of-distribution poses at test time. We show that our approach outperforms recent baselines and creates realistic animations for diverse characters. We also show that our model can handle discrete state changes, for example a profile facing left or right, that the different layers do indeed learn features specific for the respective keypoints in those layers, and that our model scales to larger datasets when more data is available. Code is available at https://github.com/tohinz/CharacterGAN. Tobias Hinz, Matthew Fisher, Oliver Wang, Eli Shechtman, Stefan Wermter |
WACV | 1 |
| 2022 | Semantic Object Accuracy for Generative Text-to-Image SynthesisabstractGenerative adversarial networks conditioned on textual image descriptions are capable of generating realistic-looking images. However, current methods still struggle to generate images based on complex image captions from a heterogeneous domain. Furthermore, quantitatively evaluating these text-to-image models is challenging, as most evaluation metrics only judge image quality but not the conformity between the image and its caption. To address these challenges we introduce a new model that explicitly models individual objects within an image and a new evaluation metric called Semantic Object Accuracy (SOA) that specifically evaluates images given an image caption. The SOA uses a pre-trained object detector to evaluate if a generated image contains objects that are mentioned in the image caption, e.g., whether an image generated from "a car driving down the street" contains a car. We perform a user study comparing several text-to-image models and show that our SOA metric ranks the models the same way as humans, whereas other metrics such as the Inception Score do not. Our evaluation also shows that models which explicitly model objects outperform models which only model global image characteristics. Tobias Hinz, Stefan Heinrich, Stefan Wermter |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | ASSET: autoregressive semantic scene editing with transformers at high resolutionsabstractWe present ASSET, a neural architecture for automatically modifying an input high-resolution image according to a user's edits on its semantic segmentation map. Our architecture is based on a transformer with a novel attention mechanism. Our key idea is to sparsify the transformer's attention matrix at high resolutions, guided by dense attention extracted at lower image resolutions. While previous attention mechanisms are computationally too expensive for handling high-resolution images or are overly constrained within specific image regions hampering long-range interactions, our novel attention mechanism is both computationally efficient and effective. Our sparsified attention mechanism is able to capture long-range interactions and context, leading to synthesizing interesting phenomena in scenes, such as reflections of landscapes onto water or fora consistent with the rest of the landscape, that were not possible to generate reliably with previous convnets and transformer approaches. We present qualitative and quantitative results, along with user studies, demonstrating the effectiveness of our method. Our code and dataset are available at our project page: https://github.com/DifanLiu/ASSET Difan Liu, Sandesh Shetty, Tobias Hinz, Matthew Fisher, Richard Zhang 0001, Taesung Park, Evangelos Kalogerakis |
ACM Trans. Graph. | 3 |
| 2021 | Encoding Complexity Analysis and Reduction for a Practically-Oriented VVC Encoder ImplementationabstractThe latest international Versatile Video Coding (VVC) standard was finalized in July 2020 by the Joint Video Experts Team (JVET) from ITU-T and ISO/IEC. Compared to the High Efficiency Video Coding (HEVC) standard, VVC offers up to 50% of bitrate savings. However, the encoder runtime of the reference software increases tenfold compared to the HEVC equivalent. This paper shows that VVC coding tools which do not consume much of the reference encoder time can have prohibitive computational cost for practical encoders, such as VVenC. VVenC is an optimized open-source real-world implementation with high relevance for practical applications. After analyzing the encoding tools in VTM and VVenC, Symmetric Motion Vector Difference (SMVD) was identified as one of the tools whose efficiency-complexity trade-off may be unfavorable for practical applications. For that reason, two complexity reduction methods for the SMVD search are proposed in this paper. The experimental evaluation on VVenC encoder at an optimized operation point confirms the SMVD encoding runtime reduction by over 72%, with minimal impact on the encoding efficiency. Ivan Zupancic, Benjamin Bross, Tobias Hinz, Detlev Marpe |
PCS | 3 |
| 2021 | Improved Techniques for Training Single-Image GANsabstractRecently there has been an interest in the potential of learning generative models from a single image, as opposed to from a large dataset. This task is of significance, as it means that generative models can be used in domains where collecting a large dataset is not feasible. However, training a model capable of generating realistic images from only a single sample is a difficult problem. In this work, we conduct a number of experiments to understand the challenges of training these methods and propose some best practices that we found allowed us to generate improved results over previous work. One key piece is that, unlike prior single image generation methods, we concurrently train several stages in a sequential multi-stage manner, allowing us to learn models with fewer stages of increasing image resolution. Compared to a recent state of the art baseline, our model is up to six times faster to train, has fewer parameters, and can better capture the global structure of images. Tobias Hinz, Matthew Fisher, Oliver Wang, Stefan Wermter |
WACV | 1 |
| 2021 | Adversarial text-to-image synthesis: A reviewabstractWith the advent of generative adversarial networks, synthesizing images from text descriptions has recently become an active research area. It is a flexible and intuitive way for conditional image generation with significant progress in the last years regarding visual realism, diversity, and semantic alignment. However, the field still faces several challenges that require further research efforts such as enabling the generation of high-resolution images with multiple objects, and developing suitable and reliable evaluation metrics that correlate with human judgement. In this review, we contextualize the state of the art of adversarial text-to-image synthesis models, their development since their inception five years ago, and propose a taxonomy based on the level of supervision. We critically examine current strategies to evaluate text-to-image synthesis models, highlight shortcomings, and identify new areas of research, ranging from the development of better datasets and evaluation metrics to possible improvements in architectural design and model training. This review complements previous surveys on generative adversarial networks with a focus on text-to-image synthesis which we believe will help researchers to further advance the field. Stanislav Frolov, Tobias Hinz, Federico Raue, Jörn Hees, Andreas Dengel 0001 |
Neural Networks | 2 |
| 2020 | Towards Fast and Efficient VVC EncodingabstractVersatile Video Coding (VVC) is a new international video coding standard to be finalized in July 2020. It is designed to provide around 50% bit-rate saving at the same subjective visual quality over its predecessor, High Efficiency Video Coding (H.265/HEVC). During the standard development, objective bit-rate savings of around 40% have been reported for the VVC reference software (VTM) compared to the HEVC reference software (HM). The unoptimized VTM encoder is around 9x, and the decoder around 2x, slower than HM. This paper discusses the VVC encoder complexity in terms of soft-ware runtime. The modular design of the standard allows a VVC encoder to trade off bit-rate savings and encoder runtime. Based on a detailed tradeoff analysis, results for different operating points are reported. Additionally, initial work on software and algorithm optimization is presented. With the optimized software algorithms, an operating point with an over 22x faster single-threaded encoder runtime than VTM can be achieved, i.e. around 2.5x faster than HM, while still providing more than 30% bit-rate savings over HM. Finally, our experiments demonstrate the flexibility of VVC and its potential for optimized soft-ware encoder implementations. Jens Brandenburg, Adam Wieckowski, Tobias Hinz, Anastasia Henkel, Valeri George, Ivan Zupancic, Christian Stoffers, Benjamin Bross, Heiko Schwarz, Detlev Marpe |
MMSP | 3 |
| 2020 | Video Compression Using Generalized Binary Partitioning, Trellis Coded Quantization, Perceptually Optimized Encoding, and Advanced Prediction and Transform CodingabstractIn this paper, we describe a video coding design that enables a higher coding efficiency than the HEVC standard. The proposed video codec follows the design of block-based hybrid video coding, but includes a number of advanced coding tools. A part of the incorporated advanced concepts was developed by the Joint Video Exploration Team, while others are newly proposed. The key aspects of these newly proposed tools are the following. A video frame is subdivided into rectangles of variable size using a binary partitioning with variable split ratios. Three new approaches for generating spatial intra prediction signals are supported: A line-wise application of conventional intra prediction modes, coupled with a mode-dependent processing order, a region-based template matching prediction method and intra prediction modes based on neural networks. For motion-compensated prediction, a multi-hypothesis mode with more than two motion hypotheses can be used. In transform coding, mode dependent combinations of primary and secondary transforms are applied. Moreover, scalar quantization is replaced by trellis-coded quantization and the entropy coding of the quantized transform coefficients is improved. The intra and inter prediction signals can be filtered using an edge-preserving diffusion filter or a non-linear DCT-based thresholding operation. The video codec includes an adaptive in-loop filter for which one of three classifiers can be chosen on a picture basis. We also incorporated an optional encoder control, which adjusts the quantization parameters based on a perceptually motivated distortion measure. In a random access scenario, our proposed video codec achieves luma BD-rate savings between 32.5% for HDR HLG UHD and 39.6% for SDR UHD over the HEVC (HM software) anchor for different categories of test sequences. Jonathan Pfaff, Heiko Schwarz, Detlev Marpe, Benjamin Bross, Santiago De-Luxán-Hernández, Philipp Helle, Christian R. Helmrich, Tobias Hinz, Wang-Q Lim, Jackie Ma, Tung Nguyen 0001, Jennifer Rasch, Michael Schäfer 0003, Mischa Siekmann, Gayathri Venugopal, Adam Wieckowski, Martin Winken, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2019 | Evaluating Defensive Distillation for Defending Text Processing Neural Networks Against Adversarial Examples
Marcus Soll, Tobias Hinz, Sven Magg, Stefan Wermter |
ICANN (3) | 2 |
| 2019 | Generating Multiple Objects at Spatially Distinct Locations
Tobias Hinz, Stefan Heinrich, Stefan Wermter |
ICLR (Poster) | 1 |
| 2018 | Inferencing based on unsupervised learning of disentangled representations
Tobias Hinz, Stefan Wermter |
ESANN | 1 |
| 2018 | Improved Prediction Via Thresholding Transform CoefficientsabstractThis paper presents a thresholding method for processing the predicted samples in the state-of-the-art High Efficiency Video Coding (HEVC) standard. The method applies an integer-based approximation of the discrete cosine transform to an extended prediction block and sets transform coefficients beneath a certain threshold to zero. Transforming back into the sample domain yields the improved prediction signal. The method is incorporated into a software implementation that is conforming to the HEVC standard and applies to both intra and inter predictions. Consequently, bit-rate savings ranging from 2.3% to 8.0% have been measured in terms of the Bjøntegaard-Delta bit rate (BD-rate). Michael Schäfer 0003, Jonathan Pfaff, Jennifer Rasch, Tobias Hinz, Heiko Schwarz, Tung Nguyen 0001, Gerhard Tech, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 4 |
| 2018 | Image Generation and Translation with Disentangled RepresentationsabstractGenerative models have made significant progress in the tasks of modeling complex data distributions such as natural images. The introduction of Generative Adversarial Networks (GANs) and auto-encoders lead to the possibility of training on big data sets in an unsupervised manner. However, for many generative models it is not possible to specify what kind of image should be generated and it is not possible to translate existing images into new images of similar domains. Furthermore, models that can perform image-to-image translation often need distinct models for each domain, making it hard to scale these systems to multiple domain image-to-image translation. We introduce a model that can do both, controllable image generation and image-to-image translation between multiple domains. We split our image representation into two parts encoding unstructured and structured information respectively. The latter is designed in a disentangled manner, so that different parts encode different image characteristics. We train an encoder to encode images into these representations and use a small amount of labeled data to specify what kind of information should be encoded in the disentangled part. A generator is trained to generate images from these representations using the characteristics provided by the disentangled part of the representation. Through this we can control what kind of images the generator generates, translate images between different domains, and even learn unknown data-generating factors while only using one single model. Tobias Hinz, Stefan Wermter |
IJCNN | 1 |
| 2018 | Speeding up the Hyperparameter Optimization of Deep Convolutional Neural NetworksabstractMost learning algorithms require the practitioner to manually set the values of many hyperparameters before the learning process can begin. However, with modern algorithms, the evaluation of a given hyperparameter setting can take a considerable amount of time and the search space is often very high-dimensional. We suggest using a lower-dimensional representation of the original data to quickly identify promising areas in the hyperparameter space. This information can then be used to initialize the optimization algorithm for the original, higher-dimensional data. We compare this approach with the standard procedure of optimizing the hyperparameters only on the original input. We perform experiments with various state-of-the-art hyperparameter optimization algorithms such as random search, the tree of parzen estimators (TPEs), sequential model-based algorithm configuration (SMAC), and a genetic algorithm (GA). Our experiments indicate that it is possible to speed up the optimization process by using lower-dimensional data representations at the beginning, while increasing the dimensionality of the input later in the optimization process. This is independent of the underlying optimization procedure, making the approach promising for many existing hyperparameter optimization algorithms. Tobias Hinz, Nicolás Navarro-Guerrero, Sven Magg, Stefan Wermter |
Int. J. Comput. Intell. Appl. | 1 |
| 2016 | The Effects of Regularization on Learning Facial Expressions with Convolutional Neural Networks
Tobias Hinz, Pablo V. A. Barros, Stefan Wermter |
ICANN (2) | 1 |
| 2013 | A Scalable Video Coding Extension of HEVCabstractThe paper describes a scalable video coding extension of the upcoming HEVC video coding standard for spatial and quality scalable coding. Besides coding tools known from scalable profiles of prior video coding standards, it includes new coding tools that further improve the enhancement layer coding efficiency. The effectiveness of the proposed scalable HEVC extension is demonstrated by comparing the coding efficiency to simulcast and single-layer coding for several test sequences and coding conditions. Philipp Helle, Haricharan Lakshman, Mischa Siekmann, Jan Stegemann, Tobias Hinz, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
DCC | 5 |
| 2013 | 3D High-Efficiency Video Coding for Multi-View Video and Depth DataabstractThis paper describes an extension of the high efficiency video coding (HEVC) standard for coding of multi-view video and depth data. In addition to the known concept of disparity-compensated prediction, inter-view motion parameter, and inter-view residual prediction for coding of the dependent video views are developed and integrated. Furthermore, for depth coding, new intra coding modes, a modified motion compensation and motion vector coding as well as the concept of motion parameter inheritance are part of the HEVC extension. A novel encoder control uses view synthesis optimization, which guarantees that high quality intermediate views can be generated based on the decoded data. The bitstream format supports the extraction of partial bitstreams, so that conventional 2D video, stereo video, and the full multi-view video plus depth format can be decoded from a single bitstream. Objective and subjective results are presented, demonstrating that the proposed approach provides 50% bit rate savings in comparison with HEVC simulcast and 20% in comparison with a straightforward multi-view extension of HEVC without the newly developed coding tools. Karsten Müller 0001, Heiko Schwarz, Detlev Marpe, Christian Bartnik, Sebastian Bosse, Heribert Brust, Tobias Hinz, Haricharan Lakshman, Philipp Merkle, Hunn Rhee, Gerhard Tech, Martin Winken, Thomas Wiegand 0001 |
IEEE Trans. Image Process. | 7 |
| 2012 | Extension of High Efficiency Video Coding (HEVC) for multiview video and depth dataabstractThis paper presents an approach for 3D video coding that uses a format in which a small number of views as well as associated depth maps are coded and transmitted. At the receiver side, additional views required for displaying the 3D video on an autostereoscopic display can be generated based on the corresponding decoded signals by using depth image based rendering (DIBR) techniques. In terms of coding technology, the proposed coding scheme represents an extension of High Efficiency Video Coding (HEVC), similar to the Multiview Coding (MVC) extension of H.264/AVC. Besides the well-known disparity-compensated prediction, advanced techniques for inter-view and inter-component prediction, the representation of depth blocks, and the encoder control for depth signals have been developed and integrated. In comparison to simulcasting the different signals using HEVC, the proposed approach provides about 40% and 50% average bit rate savings for a whole test set when configured to comply with a 2- and 3-view scenario, respectively. The proposed codec was submitted as response to a Call for Proposals on 3D Video Technology issued by the ISO/IEC Moving Picture Experts Group (MPEG) and it was ranked as the overall best performing HEVC-based proposal in the related subjective tests. Heiko Schwarz, Christian Bartnik, Sebastian Bosse, Heribert Brust, Tobias Hinz, Haricharan Lakshman, Philipp Merkle, Karsten Müller 0001, Hunn Rhee, Gerhard Tech, Martin Winken, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 5 |
| 2012 | Encoder control for renderable regions in high efficiency multiview video plus depth codingabstractThis paper describes a new encoder control method for multiview video plus depth coding. Since large parts of a multiview scenery are present in more than one of the captured video sequences, a depth-aware encoder control is introduced, which identifies those regions based on given depth maps and omits the coding of the residual signal for those regions. Experimental results indicate that bit rate reductions of about 5-9 %, depending on the bit rate, can be achieved for the 2-view case at a constant subjective quality. Sebastian Bosse, Heiko Schwarz, Tobias Hinz, Thomas Wiegand 0001 |
PCS | 3 |
| 2012 | 3D video coding using advanced prediction, depth modeling, and encoder control methodsabstractThe presented approach for 3D video coding uses the multiview video plus depth format, in which a small number of video views as well as associated depth maps are coded. Based on the coded signals, additional views required for displaying the 3D video on an autostereoscopic display can be generated by depth image based rendering techniques. The developed coding scheme represents an extension of HEVC, similar to the MVC extension of H.264/AVC. However, in addition to the well-known disparity-compensated prediction advanced techniques for inter-view and inter-component prediction, the representation of depth blocks, and the encoder control for depth signals have been integrated. In comparison to simulcasting the different signals using HEVC, the proposed approach provides about 40% and 50% bit rate savings for the tested configurations with 2 and 3 views, respectively. Bit rate reductions of about 20% have been obtained in comparison to a straightforward multiview extension of HEVC without the newly developed coding tools. Heiko Schwarz, Christian Bartnik, Sebastian Bosse, Heribert Brust, Tobias Hinz, Haricharan Lakshman, Detlev Marpe, Philipp Merkle, Karsten Müller 0001, Hunn Rhee, Gerhard Tech, Martin Winken, Thomas Wiegand 0001 |
PCS | 5 |
| 2010 | Highly efficient video compression using quadtree structures and improved techniques for motion representation and entropy codingabstractThis paper describes a novel video coding scheme that can be considered as a generalization of the block-based hybrid video coding approach of H.264/AVC. While the individual building blocks of our approach are kept simple similarly as in H.264/AVC, the flexibility of the block partitioning for prediction and transform coding has been substantially increased. This is achieved by the use of nested and pre-configurable quadtree structures, such that the block partitioning for temporal and spatial prediction as well as the space-frequency resolution of the corresponding prediction residual can be adapted to the given video signal in a highly flexible way. In addition, techniques for an improved motion representation as well as a novel entropy coding concept are included. The presented video codec was submitted to a Call for Proposals of ITU-T VCEG and ISO/IEC MPEG and was ranked among the five best performing proposals, both in terms of subjective and objective quality. Detlev Marpe, Heiko Schwarz, Sebastian Bosse, Benjamin Bross, Philipp Helle, Tobias Hinz, Heiner Kirchhoffer, Haricharan Lakshman, Tung Nguyen 0001, Simon Oudin, Mischa Siekmann, Karsten Sühring, Martin Winken, Thomas Wiegand 0001 |
PCS | 6 |
| 2010 | Video Compression Using Nested Quadtree Structures, Leaf Merging, and Improved Techniques for Motion Representation and Entropy CodingabstractAbstract-A video coding architecture is described that is based on nested and pre-configurable quadtree structures for flexible and signal-adaptive picture partitioning. The primary goal of this partitioning concept is to provide a high degree of adaptability for both temporal and spatial prediction as well as for the purpose of space-frequency representation of prediction residuals. At the same time, a leaf merging mechanism is included in order to prevent excessive partitioning of a picture into prediction blocks and to reduce the amount of bits for signaling the prediction signal. For fractional-sample motion-compensated prediction, a fixed-point implementation of the maximal-order minimum-support algorithm is presented that uses a combination of infinite impulse response and FIR filtering. Entropy coding utilizes the concept of probability interval partitioning entropy codes that offers new ways for parallelization and enhanced throughput. The presented video coding scheme was submitted to a joint call for proposals of ITU-T Visual Coding Experts Group and ISO/IEC Moving Picture Experts Group and was ranked among the five best performing proposals, both in terms of subjective and objective quality. Detlev Marpe, Heiko Schwarz, Sebastian Bosse, Benjamin Bross, Philipp Helle, Tobias Hinz, Heiner Kirchhoffer, Haricharan Lakshman, Tung Nguyen 0001, Simon Oudin, Mischa Siekmann, Karsten Sühring, Martin Winken, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2007 | Generic and Robust Video Coding with Texture Analysis and SynthesisabstractA video coding approach that uses texture analysis at the encoder and synthesis at the decoder is presented. It is characterized by scaling of texture reconstruction accuracy at the decoder according to perceptual relevance to the viewer. A closed-loop concept is proposed for a fully automated approach that can be integrated into any video codec. Besides texture analysis and synthesis, the loop encompasses a video quality assessor (VQA) that is roughly three times as complex as PSNR and performs as well as significantly more complex standardized perceptual measures. The VQA operates local and global consistency evaluations, where the outcome is utilized to correct segmentation masks provided by the texture analyzer, such that synthesis results can be improved. The proposed method is generic because rigid and non-rigid textures can be handled. Experimental results are presented that show that significant bit-rate gains can be achieved compared to H.264/AVC without our method. Patrick Ndjiki-Nya, Tobias Hinz, Thomas Wiegand 0001 |
ICME | 2 |
| 2006 | A Content-Based Video Coding Approach for Rigid and Non-Rigid TexturesabstractA generic, content-based video coding framework is described. The approach is based on H.264/MPEG4-AVC that is extended by a closed-loop texture analysis/synthesis algorithm. The texture analysis yields regions that can be reproduced with reduced accuracy at the decoder without noticeable quality degradations. These textures are synthesized at the decoder given side information generated through analysis. The remainder regions are coded using H.264/MPEG4-AVC that serves as fallback option in our framework. Texture synthesis is constrained in the proposed system as the area to synthesize is typically surrounded by one or more textures. In this paper, it is shown that constrained synthesis can be done successfully for a large class of textures including non-rigid textures such as water. Experimental results verify improvements of the proposed system compared to H.264/MPEG4-AVC without our approach. Patrick Ndjiki-Nya, Tobias Hinz, Christoph Stüber, Thomas Wiegand 0001 |
ICIP | 2 |
| 2005 | A generic and automatic content-based approach for improved H.264/MPEG4-AVC video codingabstractA new content-based approach for improved H.264/MPEG4-AVC video coding is presented. The framework is generic because it is based on a closed-loop texture analysis by synthesis algorithm that can automatically identify and recover from video quality impairments through artifact detectors and appropriate countermeasures. The algorithm is flexible, for it can in principle be integrated into any standards-compliant video codec. The fundamental assumption of our approach is that many video scenes can be classified into subjectively relevant and irrelevant textures. The texture categorization is thereby done by a texture analyzer (encoder side), while the corresponding texture synthesizer performs the replacement of the subjectively irrelevant textures (decoder side), given the side information generated by the texture analyzer. When implementing the proposed approach into an H.264/MPEG4-AVC codec, bit rate savings of up to 33.3% compared to an H.264/MPEG4-AVC video codec without our approach are reported. Patrick Ndjiki-Nya, Tobias Hinz, Aljoscha Smolic, Thomas Wiegand 0001 |
ICIP (2) | 2 |
| 2005 | Constrained inter-layer prediction for single-loop decoding in spatial scalabilityabstractThe scalability extension of H.264/AVC uses an oversampled pyramid representation for spatial scalability, where for each spatial resolution a separate motion compensation or MCTF loop is deployed. When the reconstructed signal at a lower resolution is used to predict the next higher resolution, the motion compensation or MCTF loops including the deblocking filter operations of both resolutions have to be executed. This imposes a large complexity burden on the decoding of the higher resolution signals, especially when multiple spatial layers are utilized. In this paper, we investigate the approach to only allow prediction between spatial layers for parts of the lower resolution pictures that are intra-coded in order to avoid decoding that requires multiple motion compensation or MCTF loops. Experimental results evaluate the effectiveness of the proposed approach. Heiko Schwarz, Tobias Hinz, Detlev Marpe, Thomas Wiegand 0001 |
ICIP (2) | 2 |