EDBT 2026 Demo / reviewers in the wild / expert
Fabian Brand
dblp:201/7241
· DBLP profile ↗
31ranked-venue papers
13as first author
24since 2021 · last 2025
0000-0002-2022-1033ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 13 first-author · 22 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Beyond Perspective: Neural 360-Degree Video Compression
Andy Regensky, Marc Windsheimer, Fabian Brand, André Kaup |
ICCV | 3 |
| 2025 | Boosting Neural Image Compression for Machines Using Latent Space MaskingabstractToday, many image coding scenarios do not have a human as final intended user, but rather a machine fulfilling computer vision tasks on the decoded image. Thereby, the primary goal is not to keep visual quality but maintain the task accuracy of the machine for a given bitrate. Due to the tremendous progress of deep neural networks setting benchmarking results, mostly neural networks are employed to solve the analysis tasks at the decoder side. Moreover, neural networks have also found their way into the field of image compression recently. These two developments allow for an end-to-end training of the neural compression network for an analysis network as information sink. Therefore, we first roll out such a training with a task-specific loss to enhance the coding performance of neural compression networks. Compared to the standard VVC, 41.4% of bitrate are saved by this method for Mask R-CNN as analysis network on the uncompressed Cityscapes dataset. As a main contribution, we propose LSMnet, a network that runs in parallel to the encoder network and masks out elements of the latent space that are presumably not required for the analysis network. By this approach, additional 27.3% of bitrate are saved compared to the basic neural compression network optimized with the task loss. In addition, we are the first to utilize a feature-based distortion in the training loss within the context of machine-to-machine communication, which allows for a training without annotated data. We provide extensive analyses on the Cityscapes dataset including cross-evaluation with different analysis networks and present exemplary visual results. Kristian Fischer 0001, Fabian Brand, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Overview of Variable Rate Coding in JPEG AIabstractEmpirical evidence has demonstrated that learning-based image compression can outperform classical compression frameworks. This has led to the ongoing standardization of learned-based image codecs, namely Joint Photographic Experts Group (JPEG) AI. The objective of JPEG AI is to enhance compression efficiency and provide a software and hardware-friendly solution. Based on our research, JPEG AI represents the first standardization that can facilitate the implementation of a learned image codec on a mobile device. This article presents an overview of the variable rate coding functionality in JPEG AI, which includes three variable rate adaptations: a three-dimensional quality map, a fast bit rate matching algorithm, and a training strategy. The variable rate adaptations offer a continuous rate function up to 2.0 bpp, exhibiting a high level of performance, a flexible bit allocation between different color components, and a region of interest function for the specified use case. The evaluation of performance encompasses both objective and subjective results. With regard to the objective bit rate matching, the main profile with low complexity yielded a 13.1% BD-rate gain over VVC intra, while the high profile with high complexity achieved a 19.2% BD-rate gain over VVC intra. The BD-rate result is calculated as the mean of the seven perceptual metrics defined in the JPEG AI common test conditions. With respect to subjective results, the example of improving the quality of the region of interest is illustrated. Panqi Jia, Fabian Brand, Dequan Yu, Alexander Karabutov, Elena Alshina, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Multiscale Augmented Normalizing Flows for Image CompressionabstractMost learning-based image compression methods lack efficiency for high image quality due to their non-invertible design. The decoding function of the frequently applied compressive autoencoder architecture is only an approximated inverse of the encoding transform. This issue can be resolved by using invertible latent variable models, which allow a perfect reconstruction if no quantization is performed. Furthermore, many traditional image and video coders apply dynamic block partitioning to vary the compression of certain image regions depending on their content. Inspired by this approach, hierarchical latent spaces have been applied to learning-based compression networks. In this paper, we present a novel concept, which adapts the hierarchical latent space for augmented normalizing flows, an invertible latent variable model. Our best performing model achieves significant rate savings of more than 7% over comparable single-scale models. Marc Windsheimer, Fabian Brand, André Kaup |
ICASSP | 2 |
| 2024 | ON Annotation-Free Optimization of Video Coding for MachinesabstractToday, image and video data is not only viewed by humans, but also automatically analyzed by computer vision algorithms. However, current coding standards are optimized for human perception. Emerging from this, research on video coding for machines tries to develop coding methods designed for machines as information sink. Since many of these algorithms are based on neural networks, most proposals for video coding for machines build upon neural compression. So far, optimizing the compression by applying the task loss of the analysis network, for which ground truth data is needed, is achieving the best coding performance. But ground truth data is difficult to obtain and thus an optimization without ground truth is preferred. In this paper, we present an annotation-free optimization strategy for video coding for machines. We measure the distortion by calculating the task loss of the analysis network. Therefore, the predictions on the compressed image are compared with the predictions on the original image, instead of the ground truth data. Our results show that this strategy can even outperform training with ground truth data with rate savings of up to 7.5 %. By using the non-annotated training data, the rate gains can be further increased up to 8.2 %. Marc Windsheimer, Fabian Brand, André Kaup |
ICIP | 2 |
| 2024 | Adaptive Variance-Threshold-Based Skip Modes for Learned Video Compression Using a Motion Complexity CriterionabstractSkip modes are a powerful tool to reduce the rate in video compression. The main idea is that the residual areas where the prediction performs well are not transmitted since the prediction quality is good enough that the prediction signal itself can be used as the reconstruction signal. This is commonly used, e.g., in the compression standard VVC, where a skip flag can be transmitted for inter blocks under certain conditions. The skipped residual block is then not transmitted and the content is instead inferred to be zero at the decoder. Current learning-based methods use different kinds of skip modes. One possibility here arises from the fact that the coders estimate and transmit the variance for each transmitted symbol. It has been proposed to use this estimated variance to derive a skip mode. When the variance falls below a threshold, the symbol is not transmitted. In this paper we propose an extension to this method. By classifying each position in the latent space according to the local motion complexity, we can transmit adaptive thresholds for each class. That way, we can employ motion information to refine the granularity of the skip mode. When we implement this method in FVC, we are able to save up 2.11% rate on a GOP 20 sequence. We also discuss the behavior of increasingly adaptive skip modes in scenarios with larger GOP size, where error-propagation becomes a larger issue. Fabian Brand, Jürgen Seiler, Johannes Sauer, Elena Alshina, André Kaup |
PCS | 1 |
| 2024 | Analysis of Neural Video Compression Networks for 360-Degree Video CodingabstractWith the increasing efforts of bringing high-quality virtual reality technologies into the market, efficient 360-degree video compression gains in importance. As such, the state-of-the-art H.266 NVC video coding standard integrates dedicated tools for 360-degree video, and considerable efforts have been put into designing 360-degree projection formats with improved compression efficiency. For the fast-evolving field of neural video compression networks (NVCs), the effects of different 360-degree projection formats on the overall compression performance have not yet been investigated. It is thus unclear, whether a resampling from the conventional equirectangular projection (ERP) to other projection formats yields similar gains for NVCs as for hybrid video codecs, and which formats perform best. In this paper, we analyze several generations of NVCs and an extensive set of 360-degree projection formats with respect to their compression performance for 360-degree video. Based on our analysis, we find that projection format resampling yields significant improvements in compression performance also for NVCs. The adjusted cubemap projection (ACP) and equatorial cylindrical projection (ECP) show to perform best and achieve rate savings of more than 55% compared to ERP based on WS-PSNR for the most recent NVC. Remarkably, the observed rate savings are higher than for H.266/VVC, emphasizing the importance of projection format resampling for NVCs. Andy Regensky, Fabian Brand, André Kaup |
PCS | 2 |
| 2024 | Forensic analysis of AI-compression traces in spatial and frequency domainabstractThe classical JPEG compression is a rich source of cues for forensic image analysis. However, this compression standard will in the near future be complemented by a new, highly efficient learning-based compression standard called JPEG-AI. JPEG-AI is fundamentally different from classical JPEG. Hence, its forensic traces can also be expected to be fundamentally different. We argue that there is a pressing need for image forensics research to investigate these traces. In this work, we characterize forensic compression traces of different AI compression algorithms. Our analysis investigates AI compression artifacts in frequency domain and in spatial domain. Both domains exhibit similar artifacts that likely stem from upsampling operations of the decoders. Additionally, we report for one AI codec another artifact in homogeneous regions. We also investigate the artifact detectability in several scenarios including unseen AI compression traces and postprocessing. Here, frequency and autocorrelation features are better on additive noise and classical JPEG post-compression, while RGB features perform better on blurred and downsampled images. Sandra Bergmann, Denise Moussa, Fabian Brand, André Kaup, Christian Riess |
Pattern Recognit. Lett. | 3 |
| 2024 | Conditional Residual Coding: A Remedy for Bottleneck Problems in Conditional Inter Frame CodingabstractConditional coding is a new video coding paradigm enabled by neural-network-based compression. It can be shown that conditional coding is in theory better than the traditional residual coding, which is widely used in video compression standards like HEVC or VVC. However, on closer inspection, it becomes clear that conditional coders can suffer from information bottlenecks in the prediction path, i.e., that due to the data processing inequality not all information from the prediction signal can be passed to the reconstructed signal, thereby impairing the coder performance. In this paper we propose the conditional residual coding concept, which we derive from information theoretical properties of the conditional coder. This coder significantly reduces the influence of bottlenecks, while maintaining the theoretical performance of the conditional coder. We provide a theoretical analysis of the coding paradigm and demonstrate the performance of the conditional residual coder in a practical example. We show that conditional residual coders alleviate the disadvantages of conditional coders while being able to maintain their advantages over residual coders. In the spectrum of residual and conditional coding, we can therefore consider them as “the best from both worlds”. Fabian Brand, Jürgen Seiler, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | The Bjøntegaard Bible Why Your Way of Comparing Video Codecs May Be WrongabstractIn this paper, we provide an in-depth assessment on the Bjøntegaard Delta. We construct a large data set of video compression performance comparisons using a diverse set of metrics including PSNR, VMAF, bitrate, and processing energies. These metrics are evaluated for visual data types such as classic perspective video, 360° video, point clouds, and screen content. As compression technology, we consider multiple hybrid video codecs as well as state-of-the-art neural network based compression methods. Using additional supporting points in-between standard points defined by parameters such as the quantization parameter, we assess the interpolation error of the Bjøntegaard-Delta (BD) calculus and its impact on the final BD value. From the analysis, we find that the BD calculus is most accurate in the standard application of rate-distortion comparisons with mean errors below 0.5 percentage points. For other applications and special cases, e.g., VMAF quality, energy considerations, or inter-codec comparisons, the errors are higher (up to 5 percentage points), but can be halved by using a higher number of supporting points. We finally come up with recommendations on how to use the BD calculus such that the validity of the resulting BD-values is maximized. Main recommendations are as follows: First, relative curve differences should be plotted and analyzed. Second, the logarithmic domain should be used for saturating metrics such as SSIM and VMAF. Third, BD values below a certain threshold indicated by the subset error should not be used to draw recommendations. Fourth, using two supporting points is sufficient to obtain rough performance estimates. Christian Herglotz, Hannah Och, Anna Meyer, Geetha Ramasubbu, Lena Eichermüller, Matthias Kränzler, Fabian Brand, Kristian Fischer 0001, Dat Thanh Nguyen, Andy Regensky, André Kaup |
IEEE Trans. Image Process. | 7 |
| 2023 | Saliency-Driven Hierarchical Learned Image Coding for MachinesabstractWe propose to employ a saliency-driven hierarchical neural image compression network for a machine-to-machine communication scenario following the compress-then-analyze paradigm. By that, different areas of the image are coded at different qualities depending on whether salient objects are located in the corresponding area. Areas without saliency are transmitted in latent spaces of lower spatial resolution in order to reduce the bitrate. The saliency information is explicitly derived from the detections of an object detection network. Furthermore, we propose to add saliency information to the training process in order to further specialize the different latent spaces. All in all, our hierarchical model with all proposed optimizations achieves 77.1 % bitrate savings over the latest video coding standard VVC on the Cityscapes dataset and with Mask R-CNN as analysis network at the decoder side. Thereby, it also outperforms traditional, non-hierarchical compression networks. Kristian Fischer 0001, Fabian Brand, Christian Blum 0004, André Kaup |
ICASSP | 2 |
| 2023 | Spatially-Adaptive Learning-Based Image Compression with Hierarchical Multi-Scale Latent SpacesabstractAdaptive block partitioning is responsible for large gains in current image and video compression systems. This method is able to compress large stationary image areas with only a few symbols, while maintaining a high level of quality in more detailed areas. Current state-of-the-art neural-network-based image compression systems however use only one scale to transmit the latent space. In previous publications, we proposed RDONet, a scheme to transmit the latent space in multiple spatial resolutions. Following this principle, we extend a state-of-the-art compression network by a second hierarchical latent-space level to enable multi-scale processing. We extend the existing rate variability capabilities of RDONet by a gain unit. With that we are able to outperform an equivalent traditional autoencoder by 7% rate savings. Furthermore, we show that even though we add an additional latent space, the complexity only increases marginally and the decoding time can potentially even be decreased. Fabian Brand, Alexander Kopte, Kristian Fischer 0001, André Kaup |
ICIP | 1 |
| 2023 | Processing Energy Modeling For Neural Network Based Image CompressionabstractNowadays, the compression performance of neural-network-based image compression algorithms outperforms state-of-the-art compression approaches such as JPEG or HEIC-based image compression. Unfortunately, most neural-network based compression methods are executed on GPUs and consume a high amount of energy during execution. Therefore, this paper performs an in-depth analysis on the energy consumption of state-of-the-art neural-network based compression methods on a GPU and show that the energy consumption of compression networks can be estimated using the image size with mean estimation errors of less than 7%. Finally, using a correlation analysis, we find that the number of operations per pixel is the main driving force for energy consumption and deduce that the network layers up to the second downsampling step are consuming most energy. Christian Herglotz, Fabian Brand, Andy Regensky, Felix Rievel, André Kaup |
ICIP | 2 |
| 2023 | On Interpolation of Subjective Rate-Distortion Curves for Video Coder ComparisonabstractWhen comparing two video coders, in order to obtain meaningful and unambiguous results, it is necessary to interpolate the rate-distortion curves. Typical methods like the BjØntegaard delta rate use piece-wise cubic interpolation for that purpose. This works well for the unbounded quality metric PSNR. For fitting the curve of saturating quality metrics, which are common when measuring subjective or perceptual quality, the use of a logistic fitting function was proposed with SCENIC earlier. However, the interpolation accuracy has not been validated. In this paper, we compare different methods for interpolating rate-distortion curves of subjective metrics. We come to the conclusion that SCENIC is in fact better than piece-wise cubic interpolation for four supporting points. With more supporting points, however it is restricted by the number of free parameters, and piece-wise cubic interpolation should be used. We furthermore show that logarithmic rescaling has a small benefit for the interpolation accuracy of MOS values. Fabian Brand, Christian Herglotz, André Kaup |
QoMEX | 1 |
| 2022 | Learning True Rate-Distortion-Optimization for End-To-End Image CompressionabstractEven though rate-distortion optimization is a crucial part of traditional image and video compression, not many approaches exist which transfer this concept to end-to-end-trained image compression. Most frameworks contain static compression and decompression models which are fixed after training, so efficient rate-distortion optimization is not possible. In a previous work, we proposed RDONet [1], which enables an RDO approach comparable to adaptive block partitioning in HEVC. In this paper, we enhance the training and boost the model performance by introducing low-complexity estimations of the RDO result into the training. It is well known that the setup during the training should be as close as possible to the setup during inference. Since including an RDO search into the training is computationally not feasible, we propose a fast variance-based criterion which we can use to approximate the RDO behavior during training. Additionally, we use the same criterion to propose a variance-adaptive RDO initialization which converges faster, needs fewer RDO passes. We can therefore decrease the inference runtime significantly. With our novel training method, we achieve average Bjøntegaard rate savings of 19.6% in MS-SSIM over the previous RDONet model [1], which equals rate savings of 27.3% over a comparable conventional deep image coder, similar to [2]. With our novel initialization method, we can reduce the number of RDO passes to one. Therefore, we need only half the time for RDO, while still saving 26.8% rate. When we do not perform an RDO search but instead only rely on the initial estimation, we still obtain remarkable rate-savings of 23.6%, needing no additional time for an RDO search. The full paper is available on arXiv [3]. Fabian Brand, Kristian Fischer 0001, Alexander Kopte, André Kaup |
DCC | 1 |
| 2022 | A Low-Parametric Model for Bit-Rate Estimation of VVC Residual CodingabstractThere are many tasks within video compression which re-quire fast bit rate estimation. As an example, rate-control algorithms are only feasible because it is possible to estimate the required bit rate without needing to encode the en-tire block. With residual coding technology becoming more and more sophisticated, the corresponding bit rate models re-quire more advanced features. In this work, we propose a set of four features together with a linear model, which is able to estimate the rate of arbitrary residual blocks which were compressed using the VVC standard. Our method out-performs other methods which were used for the same task both in terms of mean absolute error and mean relative error. Our model deviates by less than 4 bit on average over a large dataset of natural images. Fabian Brand, Christian Herglotz, André Kaup |
ICASSP | 1 |
| 2022 | P-Frame Coding with Generalized Difference: A Novel Conditional Coding ApproachabstractMotion compensated inter frame prediction is a common component of all video coders and greatly reduces temporal redundancy. With the rise of deep learning-based image and video compression, this concept has been successfully taken over from traditional coding approaches. These approaches offer a larger flexibility than traditional transform coding and therefore enable efficient conditional coding. In this work, we develop a novel conditional coding approach based on the generalized difference and generalized sum operators. This approach is a special case of a general conditional coder and has a very small complexity overhead. We also propose an extension which enables dynamic content-adaptive switching between conditional and residual coding. We show that the extended generalized difference coding outperforms both residual and conditional coding, saving 27.8% Bjøntegaard delta rate compared to the former. Fabian Brand, Jürgen Seiler, André Kaup |
ICIP | 1 |
| 2022 | Learning Frequency-Specific Quantization Scaling in VVC for Standard-Compliant Task-Driven Image CodingabstractToday, visual data is often analyzed by a neural network without any human being involved, which demands for specialized codecs. For standard-compliant codec adaptations towards certain information sinks, HEVC or VVC provide the possibility of frequency-specific quantization with scaling lists. This is a well-known method for the human visual system, where scaling lists are derived from psycho-visual models. In this work, we employ scaling lists when performing VVC intra coding for neural networks as information sink. To this end, we propose a novel data-driven method to obtain optimal scaling lists for arbitrary neural networks. Experiments with Mask R-CNN as information sink reveal that coding the Cityscapes dataset with the proposed scaling lists result in peak bitrate savings of 8.9 % over VVC with constant quantization. By that, our approach also outperforms scaling lists optimized for the human visual system. The generated scaling lists can be found under https://github.com/FAU-LMS/VCM_scaling_lists. Kristian Fischer 0001, Fabian Brand, Christian Herglotz, André Kaup |
ICIP | 2 |
| 2022 | Domain Adaptation for Unknown Image Distortions in Instance SegmentationabstractData-driven techniques for machine vision heavily depend on the training data to sufficiently resemble the data occurring during test and application. However, in practice unknown distortion can lead to a domain gap between training and test data, impeding the performance of a machine vision system. With our proposed approach this domain gap can be closed by unpaired learning of the pristine-to-distortion mapping function of the unknown distortion. This learned mapping function may then be used to emulate the unknown distortion in the training data. Employing a fixed setup, our approach is independent from prior knowledge of the distortion. Within this work, we show that we can effectively learn unknown distortions at arbitrary strengths. When applying our approach to instance segmentation in an autonomous driving scenario, we achieve results comparable to an oracle with knowledge of the distortion. An average gain in mean Average Precision (mAP) of up to 0.19 can be achieved. Maximiliane Gruber, Fabian Brand, Alina Mosebach, Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2022 | On Benefits and Challenges of Conditional Interframe Video Coding in Light of Information TheoryabstractThe rise of variational autoencoders for image and video compression has opened the door to many elaborate coding techniques. One example here is the possibility of conditional interframe coding. Here, instead of transmitting the residual between the original frame and the predicted frame (often obtained by motion compensation), the current frame is transmitted under the condition of knowing the prediction signal. In practice, conditional coding can be straightforwardly implemented using a conditional autoencoder, which has also shown good results in recent works. In this paper, we provide an information theoretical analysis of conditional coding for inter frames and show in which cases gains compared to traditional residual coding can be expected. We also show the effect of information bottlenecks which can occur in practical video coders in the prediction signal path due to the network structure, as a consequence of the data-processing theorem or due to quantization. We demonstrate that conditional coding has theoretical benefits over residual coding but that there are cases in which the benefits are quickly canceled by small information bottlenecks of the prediction signal. Fabian Brand, Jürgen Seiler, André Kaup |
PCS | 1 |
| 2021 | Intra To Inter: Towards Intra Prediction for Learning-Based Video Coders Using Optical FlowabstractTraditional video coders often rely on a block structure for transmission. Here each block is coded separately and sequentially and for each block the encoder can decide whether to use intra or inter prediction. This way, inter and intra prediction can be mixed within a single frame. This has advantages when new areas are uncovered, which were not present in the reference frame, and can hence not be predicted well. These areas are typically predicted using intra prediction. Currently much research goes into end-to-end-trained video coders which do not operate on a block level and typically use dense motion fields for inter prediction. There it is more difficult to incorporate intra prediction for uncovered regions. In this paper we propose a novel concept which enables us to reinterpret classical angular intra prediction in a way that we can transmit it as part of the dense motion field. We can save an average of 18% rate for the transmission of the motion vectors for the same quality of the prediction image. Fabian Brand, Jürgen Seiler, André Kaup |
ICIP | 1 |
| 2021 | A Novel End-To-End Network for Reconstruction of Non-Regularly Sampled Image Data Using Locally Fully Connected LayersabstractQuarter sampling and three-quarter sampling are novel sensor concepts that enable the acquisition of higher resolution images without increasing the number of pixels. This is achieved by non-regularly covering parts of each pixel of a low-resolution sensor such that only one quadrant or three quadrants of the sensor area of each pixel is sensitive to light. Combining a properly designed mask and a high-quality reconstruction algorithm, a higher image quality can be achieved than using a low-resolution sensor and subsequent upsampling. For the latter case, the image quality can be further enhanced using super resolution algorithms such as the very deep super resolution network (VDSR). In this paper, we propose a novel end-to-end neural network to reconstruct high resolution images from non-regularly sampled sensor data. The network is a concatenation of a locally fully connected reconstruction network (LFCR) and a standard VDSR network. Altogether, using a three-quarter sampling sensor with our novel neural network layout, the image quality in terms of PSNR for the Urban100 dataset can be increased by 2.96 dB compared to the state-of-the-art approach. Compared to a low-resolution sensor with VDSR, a gain of 1.11 dB is achieved. Simon Grosche, Fabian Brand, André Kaup |
MMSP | 2 |
| 2021 | Switchable Motion Models for Non-Block-Based Inter Prediction in Learning-Based Video CodingabstractMost state-of-the-art video coders rely on a block structure. For inter-frame prediction, motion vectors are transmitted per block. For example in VVC, the coder can choose between a translational or an affine motion model on a block level, depending on the content. In non-block-based coding, which is on the rise since the development of end-to-end learning based image compression, the motion vectors have to be transmitted differently. Due to the missing inherent block structure, switching between different motion models presents a challenge, but also an opportunity. In this paper, we propose an alternative approach to efficiently signal additional information regarding the motion model to improve the quality of the motion compensated image. Using our methods, we are able to increase the quality of the prediction image in our scenario by 0.40 dB on average and by up to 0.88 dB for sequences with strong and complex motion at the same rate. Fabian Brand, Jürgen Seiler, André Kaup |
PCS | 1 |
| 2021 | DeepCNV: a deep learning approach for authenticating copy number variationsabstractCopy number variations (CNVs) are an important class of variations contributing to the pathogenesis of many disease phenotypes. Detecting CNVs from genomic data remains difficult, and the most currently applied methods suffer from an unacceptably high false positive rate. A common practice is to have human experts manually review original CNV calls for filtering false positives before further downstream analysis or experimental validation. Here, we propose DeepCNV, a deep learning-based tool, intended to replace human experts when validating CNV calls, focusing on the calls made by one of the most accurate CNV callers, PennCNV. The sophistication of the deep neural network algorithm is enriched with over 10 000 expert-scored samples that are split into training and testing sets. Variant confidence, especially for CNVs, is a main roadblock impeding the progress of linking CNVs with the disease. We show that DeepCNV adds to the confidence of the CNV calls with an optimal area under the receiver operating characteristic curve of 0.909, exceeding other machine learning methods. The superiority of DeepCNV was also benchmarked and confirmed using an experimental wet-lab validation dataset. We conclude that the improvement obtained by DeepCNV results in significantly fewer false positive results and failures to replicate the CNV association results. Joseph Glessner, Xiurui Hou, Jie Zhang 0049, Munir Khan, Fabian Brand, Peter M. Krawitz, Patrick Sleiman, Hakon Hakonarson, Zhi Wei 0001 |
Briefings Bioinform. | 6 |
| 2020 | Enhanced Image Reconstruction From Quarter Sampling Measurements Using An Adapted Very Deep Super Resolution NetworkabstractQuarter sampling is a novel sensor concept that enables the acquisition of higher resolution images without increasing the number of pixels. This is achieved by covering three quarters of each pixel of a low-resolution sensor such that only one quadrant of the sensor area of each pixel is sensitive to light. By randomly masking different parts, effectively a non-regular sampling of a higher resolution image is performed. Combining a properly designed mask and a high-quality reconstruction algorithm, a higher image quality can be achieved than using a low-resolution sensor and subsequent upsampling. For the latter case, the image quality can be enhanced using super resolution algorithms. Recently, algorithms based on machine learning such as the Very Deep Super Resolution network (VDSR) proofed to be successful for this task. In this work, we transfer the concepts of VDSR to the special case of quarter sampling. Besides adapting the network layout to take advantage of the case of quarter sampling, we introduce a novel data augmentation technique enabled by quarter sampling. Altogether, using the quarter sampling sensor, the image quality in terms of PSNR can be increased by + 0.67 dB for the Urban 100 dataset compared to using a low-resolution sensor with VDSR. Simon Grosche, Kristian Fischer 0001, Fabian Brand, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2020 | A Triangulation-Based Backward Adaptive Motion Field Subsampling SchemeabstractOptical flow procedures are used to generate dense motion fields which approximate true motion. Such fields contain a large amount of data and if we need to transmit such a field, the raw data usually exceeds the raw data of the two images it was computed from. In many scenarios, however, it is of interest to transmit a dense motion field efficiently. Most prominently this is the case in inter prediction for video coding. In this paper we propose a transmission scheme based on subsampling the motion field. Since a field which was subsampled with a regularly spaced pattern usually yields suboptimal results, we propose an adaptive subsampling algorithm that preferably samples vectors at positions where changes in motion occur. The subsampling pattern is fully reconstructable without the need for signaling of position information. We show an average gain of 2.95 dB in average end point error compared to regular subsampling. Furthermore we show that an additional prediction stage can improve the results by an additional 0.43 dB, gaining 3.38 dB in total. Fabian Brand, Jürgen Seiler, Elena Alshina, André Kaup |
MMSP | 1 |
| 2020 | Video Coding for Machines with Feature-Based Rate-Distortion OptimizationabstractCommon state-of-the-art video codecs are optimized to deliver a low bitrate by providing a certain quality for the final human observer, which is achieved by rate-distortion optimization (RDO). But, with the steady improvement of neural networks solving computer vision tasks, more and more multimedia data is not observed by humans anymore, but directly analyzed by neural networks. In this paper, we propose a standard-compliant feature-based RDO (FRDO) that is designed to increase the coding performance, when the decoded frame is analyzed by a neural network in a video coding for machine scenario. To that extent, we replace the pixel-based distortion metrics in conventional RDO of VTM-8.0 with distortion metrics calculated in the feature space created by the first layers of a neural network. Throughout several tests with the segmentation network Mask R-CNN and single images from the Cityscapes dataset, we compare the proposed FRDO and its hybrid version HFRDO with different distortion measures in the feature space against the conventional RDO. With HFRDO, up to 5.49% bitrate can be saved compared to the VTM-8.0 implementation in terms of Bjøntegaard Delta Rate and using the weighted average precision as quality metric. Additionally, allowing the encoder to vary the quantization parameter results in coding gains for the proposed HFRDO of up 9.95% compared to conventional VTM. Kristian Fischer 0001, Fabian Brand, Christian Herglotz, André Kaup |
MMSP | 2 |
| 2020 | Introducing Latent Space Correlation to Conditional Autoencoders for Intra PredictionabstractIntra prediction has been an integral part of image and video coders for a long time. A predominant method is angular prediction that extends the reference area in a certain angle into the block. Recently many deep-learning-based methods have been proposed. Since intra prediction uses multiple modes this usually requires training a large number of networks. With a conditional autoencoder we are able to generate an arbitrary number of modes with only one network. In this paper we introduce a novel loss function enforcing a spatially correlated latent space and extend the network structure to the same end. Thereby we are able to propose a simple spatial mode prediction scheme using most-probable-mode lists. By replacing matrix-based intra prediction in VVC with our method, we obtain average rate savings of 0.84% with peak gains of 2.37%. Fabian Brand, Jürgen Seiler, André Kaup |
VCIP | 1 |
| 2019 | Mid-level Chord Transition Features for Musical Style AnalysisabstractChords and their progressions are an important aspect of Western tonal music. Specifically, transitions between subsequent chords within a piece carry style-relevant information. To extract such information from audio recordings, a naive approach first peforms automatic chord estimation for computing chord labels explicitly and then derives transition statistics. Often, this is done with Hidden Markov Models involving the Viterbi decoding algorithm. However, since chords are often ambiguous, deciding on one “optimal” chord sequence can be problematic, which heavily affects the subsequent derivation of transition features. In this paper, we propose novel mid-level features that capture chord transitions in a “soft” way. Our method exploits the Baum-Welch algorithm, which does not involve hard decisions on chord labels. Instead, we obtain probabilistic features that account for ambiguities among chords and chord transitions. In several experiments, we evaluate these features within a style classification scenario discriminating four historical periods of Western classical music. Our soft transition features consistently achieve higher accuracies than comparable hard-decision features, thus demonstrating the descriptive power of the novel features. Christof Weiß, Fabian Brand, Meinard Müller |
ICASSP | 2 |
| 2019 | Intra Frame Prediction for Video Coding Using a Conditional Autoencoder ApproachabstractIntra prediction is a vital component of most modern image and video codecs. State of the art video codecs like High Efficiency Video Coding (HEVC) or the upcoming Versatile Video Coding (VVC) use a high number of directional modes. With the recent advances in deep learning, it is now possible to use artificial neural networks for intra frame prediction. Previously published approaches usually add additional ANN based modes or replace all modes by training several networks. In our approach, we use a single autoencoder network to first compress the original with help of already transmitted pixels to four parameters. We then use the parameters together with this support area to generate a prediction for the block. This way, we are able to replace all angular intra modes by a single ANN. In the experiments we compare our method with the intra prediction method currently used in the VVC Test Model (VTM). Using our method, we are able to gain up to 0.85 dB prediction PSNR with a comparable amount of side information or reduce the amount of side information by 2 bit per prediction unit with similar PSNR. Fabian Brand, Jürgen Seiler, André Kaup |
PCS | 1 |
| 2017 | Motion compensated frame rate up-conversion using 3D frequency selective extrapolation and a multi-layer consistency checkabstractA high temporal resolution is desirable in many applications such as entertainment systems, automotive systems, or video surveillance. Apart from using cameras with a higher temporal resolution, it is also possible to employ frame rate up-conversion methods to obtain an enhanced temporal resolution. In principle, those algorithms can be grouped into approaches that rely on a motion estimation and approaches that do not. Both strategies typically process a video sequence frame by frame and take into account only the directly adjacent frames to compute the intermediate frame. In this paper, we propose a frame rate up-conversion technique that employs a motion compensated three-dimensional reconstruction algorithm. As a result, the proposed method takes into account more than two frames and is capable of jointly reconstructing up to a certain amount of missing frames in a video sequence. Furthermore, we present a multi-layer consistency check to further improve the reconstruction. On average, simulation results show a luminance PSNR gain compared to a conventional frame rate up-conversion method of 0.5 dB. Visual examples substantiate our objective results. Michel Bätz, Fabian Brand, Andrea Eichenseer, André Kaup |
ICASSP | 2 |