EDBT 2026 Demo / reviewers in the wild / expert
Kristian Fischer 0001
dblp:71/2777-1
· DBLP profile ↗
15ranked-venue papers
10as first author
10since 2021 · last 2025
0000-0002-0024-3171ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Boosting Neural Image Compression for Machines Using Latent Space MaskingabstractToday, many image coding scenarios do not have a human as final intended user, but rather a machine fulfilling computer vision tasks on the decoded image. Thereby, the primary goal is not to keep visual quality but maintain the task accuracy of the machine for a given bitrate. Due to the tremendous progress of deep neural networks setting benchmarking results, mostly neural networks are employed to solve the analysis tasks at the decoder side. Moreover, neural networks have also found their way into the field of image compression recently. These two developments allow for an end-to-end training of the neural compression network for an analysis network as information sink. Therefore, we first roll out such a training with a task-specific loss to enhance the coding performance of neural compression networks. Compared to the standard VVC, 41.4% of bitrate are saved by this method for Mask R-CNN as analysis network on the uncompressed Cityscapes dataset. As a main contribution, we propose LSMnet, a network that runs in parallel to the encoder network and masks out elements of the latent space that are presumably not required for the analysis network. By this approach, additional 27.3% of bitrate are saved compared to the basic neural compression network optimized with the task loss. In addition, we are the first to utilize a feature-based distortion in the training loss within the context of machine-to-machine communication, which allows for a training without annotated data. We provide extensive analyses on the Cityscapes dataset including cross-evaluation with different analysis networks and present exemplary visual results. Kristian Fischer 0001, Fabian Brand, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | The Bjøntegaard Bible Why Your Way of Comparing Video Codecs May Be WrongabstractIn this paper, we provide an in-depth assessment on the Bjøntegaard Delta. We construct a large data set of video compression performance comparisons using a diverse set of metrics including PSNR, VMAF, bitrate, and processing energies. These metrics are evaluated for visual data types such as classic perspective video, 360° video, point clouds, and screen content. As compression technology, we consider multiple hybrid video codecs as well as state-of-the-art neural network based compression methods. Using additional supporting points in-between standard points defined by parameters such as the quantization parameter, we assess the interpolation error of the Bjøntegaard-Delta (BD) calculus and its impact on the final BD value. From the analysis, we find that the BD calculus is most accurate in the standard application of rate-distortion comparisons with mean errors below 0.5 percentage points. For other applications and special cases, e.g., VMAF quality, energy considerations, or inter-codec comparisons, the errors are higher (up to 5 percentage points), but can be halved by using a higher number of supporting points. We finally come up with recommendations on how to use the BD calculus such that the validity of the resulting BD-values is maximized. Main recommendations are as follows: First, relative curve differences should be plotted and analyzed. Second, the logarithmic domain should be used for saturating metrics such as SSIM and VMAF. Third, BD values below a certain threshold indicated by the subset error should not be used to draw recommendations. Fourth, using two supporting points is sufficient to obtain rough performance estimates. Christian Herglotz, Hannah Och, Anna Meyer, Geetha Ramasubbu, Lena Eichermüller, Matthias Kränzler, Fabian Brand, Kristian Fischer 0001, Dat Thanh Nguyen, Andy Regensky, André Kaup |
IEEE Trans. Image Process. | 8 |
| 2023 | Saliency-Driven Hierarchical Learned Image Coding for MachinesabstractWe propose to employ a saliency-driven hierarchical neural image compression network for a machine-to-machine communication scenario following the compress-then-analyze paradigm. By that, different areas of the image are coded at different qualities depending on whether salient objects are located in the corresponding area. Areas without saliency are transmitted in latent spaces of lower spatial resolution in order to reduce the bitrate. The saliency information is explicitly derived from the detections of an object detection network. Furthermore, we propose to add saliency information to the training process in order to further specialize the different latent spaces. All in all, our hierarchical model with all proposed optimizations achieves 77.1 % bitrate savings over the latest video coding standard VVC on the Cityscapes dataset and with Mask R-CNN as analysis network at the decoder side. Thereby, it also outperforms traditional, non-hierarchical compression networks. Kristian Fischer 0001, Fabian Brand, Christian Blum 0004, André Kaup |
ICASSP | 1 |
| 2023 | Spatially-Adaptive Learning-Based Image Compression with Hierarchical Multi-Scale Latent SpacesabstractAdaptive block partitioning is responsible for large gains in current image and video compression systems. This method is able to compress large stationary image areas with only a few symbols, while maintaining a high level of quality in more detailed areas. Current state-of-the-art neural-network-based image compression systems however use only one scale to transmit the latent space. In previous publications, we proposed RDONet, a scheme to transmit the latent space in multiple spatial resolutions. Following this principle, we extend a state-of-the-art compression network by a second hierarchical latent-space level to enable multi-scale processing. We extend the existing rate variability capabilities of RDONet by a gain unit. With that we are able to outperform an equivalent traditional autoencoder by 7% rate savings. Furthermore, we show that even though we add an additional latent space, the complexity only increases marginally and the decoding time can potentially even be decreased. Fabian Brand, Alexander Kopte, Kristian Fischer 0001, André Kaup |
ICIP | 3 |
| 2022 | Learning True Rate-Distortion-Optimization for End-To-End Image CompressionabstractEven though rate-distortion optimization is a crucial part of traditional image and video compression, not many approaches exist which transfer this concept to end-to-end-trained image compression. Most frameworks contain static compression and decompression models which are fixed after training, so efficient rate-distortion optimization is not possible. In a previous work, we proposed RDONet [1], which enables an RDO approach comparable to adaptive block partitioning in HEVC. In this paper, we enhance the training and boost the model performance by introducing low-complexity estimations of the RDO result into the training. It is well known that the setup during the training should be as close as possible to the setup during inference. Since including an RDO search into the training is computationally not feasible, we propose a fast variance-based criterion which we can use to approximate the RDO behavior during training. Additionally, we use the same criterion to propose a variance-adaptive RDO initialization which converges faster, needs fewer RDO passes. We can therefore decrease the inference runtime significantly. With our novel training method, we achieve average Bjøntegaard rate savings of 19.6% in MS-SSIM over the previous RDONet model [1], which equals rate savings of 27.3% over a comparable conventional deep image coder, similar to [2]. With our novel initialization method, we can reduce the number of RDO passes to one. Therefore, we need only half the time for RDO, while still saving 26.8% rate. When we do not perform an RDO search but instead only rely on the initial estimation, we still obtain remarkable rate-savings of 23.6%, needing no additional time for an RDO search. The full paper is available on arXiv [3]. Fabian Brand, Kristian Fischer 0001, Alexander Kopte, André Kaup |
DCC | 2 |
| 2022 | Evaluation of Video Coding for Machines without Ground TruthabstractIn the emerging field of video coding for machines, video datasets with pristine video quality and high-quality annotations are required for a comprehensive evaluation. However, existing video datasets with detailed annotations are severely limited in size and video quality. Thus, current methods have to either evaluate their codecs on still images or on already compressed data. To mitigate this problem, we propose an evaluation method based on pseudo ground-truth data from the field of semantic segmentation to the evaluation of video coding for machines. Through extensive evaluation, this paper shows that the proposed ground-truth-agnostic evaluation method results in an acceptable absolute measurement error below 0.7 percentage points on the Bjøntegaard Delta Rate compared to using the true ground truth for mid-range bitrates. We evaluate on the three tasks of semantic segmentation, instance segmentation, and object detection. Lastly, we utilize the ground-truth-agnostic method to measure the coding performances of the VVC compared against HEVC on the Cityscapes sequences. This reveals that the coding position has a significant influence on the task performance. Kristian Fischer 0001, Markus Hofbauer, Christopher B. Kuhn, Eckehard G. Steinbach, André Kaup |
ICASSP | 1 |
| 2022 | Learning Frequency-Specific Quantization Scaling in VVC for Standard-Compliant Task-Driven Image CodingabstractToday, visual data is often analyzed by a neural network without any human being involved, which demands for specialized codecs. For standard-compliant codec adaptations towards certain information sinks, HEVC or VVC provide the possibility of frequency-specific quantization with scaling lists. This is a well-known method for the human visual system, where scaling lists are derived from psycho-visual models. In this work, we employ scaling lists when performing VVC intra coding for neural networks as information sink. To this end, we propose a novel data-driven method to obtain optimal scaling lists for arbitrary neural networks. Experiments with Mask R-CNN as information sink reveal that coding the Cityscapes dataset with the proposed scaling lists result in peak bitrate savings of 8.9 % over VVC with constant quantization. By that, our approach also outperforms scaling lists optimized for the human visual system. The generated scaling lists can be found under https://github.com/FAU-LMS/VCM_scaling_lists. Kristian Fischer 0001, Fabian Brand, Christian Herglotz, André Kaup |
ICIP | 1 |
| 2021 | Saliency-Driven Versatile Video Coding for Neural Object DetectionabstractSaliency-driven image and video coding for humans has gained importance in the recent past. In this paper, we pro-pose such a saliency-driven coding framework for the video coding for machines task using the latest video coding standard Versatile Video Coding (VVC). To determine the salient regions before encoding, we employ the real-time-capable object detection network You Only Look Once (YOLO) in combination with a novel decision criterion. To measure the coding quality for a machine, the state-of-the-art object segmentation network Mask R-CNN was applied to the decoded frame. From extensive simulations we find that, compared to the reference VVC with a constant quality, up to 29 % of bitrate can be saved with the same detection accuracy at the decoder side by applying the proposed saliency-driven framework. Besides, we compare YOLO against other, more traditional saliency detection methods. Kristian Fischer 0001, Felix Fleckenstein, Christian Herglotz, André Kaup |
ICASSP | 1 |
| 2021 | Analysis Of Neural Image Compression Networks For Machine-To-Machine CommunicationabstractVideo and image coding for machines (VCM) is an emerging field that aims to develop compression methods resulting in optimal bitstreams when the decoded frames are analyzed by a neural network. Several approaches already exist improving classic hybrid codecs for this task. However, neural compression networks (NCNs) have made an enormous progress in coding images over the last years. Thus, it is reasonable to consider such NCNs, when the information sink at the decoder side is a neural network as well. Therefore, we build-up an evaluation framework analyzing the performance of four state-of-the-art NCNs, when a Mask R-CNN is segmenting objects from the decoded image. The compression performance is measured by the weighted average precision for the Cityscapes dataset. Based on that analysis, we find that networks with leaky ReLU as non-linearity and training with SSIM as distortion criteria results in the highest coding gains for the VCM task. Furthermore, it is shown that the GAN-based NCN architecture achieves the best coding performance and even out-performs the recently standardized Versatile Video Coding (VVC) for the given scenario. Kristian Fischer 0001, Christian Forsch, Christian Herglotz, André Kaup |
ICIP | 1 |
| 2021 | Robust Deep Neural Object Detection and Segmentation for Automotive Driving Scenario with Compressed Image DataabstractDeep neural object detection or segmentation networks are commonly trained with pristine, uncompressed data. However, in practical applications the input images are usually deteriorated by compression that is applied to efficiently transmit the data. Thus, we propose to add deteriorated images to the training process in order to increase the robustness of the two state-of-the-art networks Faster and Mask R-CNN. Throughout our paper, we investigate an autonomous driving scenario by evaluating the newly trained models on the Cityscapes dataset that has been compressed with the upcoming video coding standard Versatile Video Coding (VVC). When employing the models that have been trained with the proposed method, the weighted average precision of the R-CNNs can be increased by up to 3.68 percentage points for compressed input images, which corresponds to bitrate savings of nearly 48 %. Kristian Fischer 0001, Christian Blum 0004, Christian Herglotz, André Kaup |
ISCAS | 1 |
| 2020 | On Intra Video Coding And In-Loop Filtering For Neural Object Detection NetworksabstractClassical video coding for satisfying humans as the final user is a widely investigated field of studies for visual content, and common video codecs are all optimized for the human visual system (HVS). But are the assumptions and optimizations also valid when the compressed video stream is analyzed by a machine? To answer this question, we compared the performance of two state-of-the-art neural detection networks when being fed with deteriorated input images coded with HEVC and VVC in an autonomous driving scenario using intra coding. Additionally, the impact of the three VVC in-loop filters when coding images for a neural network is examined. The results are compared using the mean average precision metric to evaluate the object detection performance for the compressed inputs. Throughout these tests, we found that the Bjøntegaard Delta Rate savings with respect to PSNR of 22.2 % using VVC instead of HEVC cannot be reached when coding for object detection networks with only 13.6 % in the best case. Besides, it is shown that disabling the VVC in-loop filters SAO and ALF results in bitrate savings of 6.4 % compared to the standard VTM at the same mean average precision. Kristian Fischer 0001, Christian Herglotz, André Kaup |
ICIP | 1 |
| 2020 | Enhanced Image Reconstruction From Quarter Sampling Measurements Using An Adapted Very Deep Super Resolution NetworkabstractQuarter sampling is a novel sensor concept that enables the acquisition of higher resolution images without increasing the number of pixels. This is achieved by covering three quarters of each pixel of a low-resolution sensor such that only one quadrant of the sensor area of each pixel is sensitive to light. By randomly masking different parts, effectively a non-regular sampling of a higher resolution image is performed. Combining a properly designed mask and a high-quality reconstruction algorithm, a higher image quality can be achieved than using a low-resolution sensor and subsequent upsampling. For the latter case, the image quality can be enhanced using super resolution algorithms. Recently, algorithms based on machine learning such as the Very Deep Super Resolution network (VDSR) proofed to be successful for this task. In this work, we transfer the concepts of VDSR to the special case of quarter sampling. Besides adapting the network layout to take advantage of the case of quarter sampling, we introduce a novel data augmentation technique enabled by quarter sampling. Altogether, using the quarter sampling sensor, the image quality in terms of PSNR can be increased by + 0.67 dB for the Urban 100 dataset compared to using a low-resolution sensor with VDSR. Simon Grosche, Kristian Fischer 0001, Fabian Brand, Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2020 | Video Coding for Machines with Feature-Based Rate-Distortion OptimizationabstractCommon state-of-the-art video codecs are optimized to deliver a low bitrate by providing a certain quality for the final human observer, which is achieved by rate-distortion optimization (RDO). But, with the steady improvement of neural networks solving computer vision tasks, more and more multimedia data is not observed by humans anymore, but directly analyzed by neural networks. In this paper, we propose a standard-compliant feature-based RDO (FRDO) that is designed to increase the coding performance, when the decoded frame is analyzed by a neural network in a video coding for machine scenario. To that extent, we replace the pixel-based distortion metrics in conventional RDO of VTM-8.0 with distortion metrics calculated in the feature space created by the first layers of a neural network. Throughout several tests with the segmentation network Mask R-CNN and single images from the Cityscapes dataset, we compare the proposed FRDO and its hybrid version HFRDO with different distortion measures in the feature space against the conventional RDO. With HFRDO, up to 5.49% bitrate can be saved compared to the VTM-8.0 implementation in terms of Bjøntegaard Delta Rate and using the weighted average precision as quality metric. Additionally, allowing the encoder to vary the quantization parameter results in coding gains for the proposed HFRDO of up 9.95% compared to conventional VTM. Kristian Fischer 0001, Fabian Brand, Christian Herglotz, André Kaup |
MMSP | 1 |
| 2020 | Decoding-Energy Optimal Video Encoding For x265abstractThis paper presents optimal x265-encoder configurations and an enhanced optimization algorithm for minimizing the software decoding energy of HEVC-coded videos. We reach this goal with two contributions. First, we perform a detailed analysis on the influence of various encoder settings on the decoding energy. Second, we include an enhanced version of an algorithm called decoding-energy-rate-distortion optimization into x265, which we optimize for fast and efficient encoding. This algorithm introduces the estimated decoding energy as an additional optimization criterion into the rate-distortion cost function. We evaluate the extended encoder in terms of bitrate, distortion, and decoding energy, where we perform energy measurements to prove the superior energy efficiency. We find that the combination of the `fastdecoding' tuning option of x265 with the enhanced decoding-energy-rate-distortion optimization leads to 27.2% and 26.0% of decoding energy savings for OpenHEVC and HM decoding, respectively. At the same time, compression efficiency losses of 38.2% and negligible decreases in encoder runtime of 0.39% can be observed. Christian Herglotz, Marco Bader, Kristian Fischer 0001, André Kaup |
MMSP | 3 |
| 2020 | On Versatile Video Coding at UHD with Machine-Learning-Based Super-ResolutionabstractCoding 4K data has become of vital interest in recent years, since the amount of 4K data is significantly increasing. We propose a coding chain with spatial down- and upscaling that combines the next-generation VVC codec with machine learning based single image super-resolution algorithms for 4K. The investigated coding chain, which spatially downscales the 4K data before coding, shows superior quality than the conventional VVC reference software for low bitrate scenarios. Throughout several tests, we find that up to 12 % and 18% Bj⊘ntegaard delta rate gains can be achieved on average when coding 4K sequences with VVC and QP values above 34 and 42, respectively. Additionally, the investigated scenario with up- and downscaling helps to reduce the loss of details and compression artifacts, as it is shown in a visual example. Kristian Fischer 0001, Christian Herglotz, André Kaup |
QoMEX | 1 |