Christian Rohlfing

dblp:120/8288 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
9since 2021 · last 2023
0000-0002-0372-056XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 8 since 2021Security and privacy · 1 · 1 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Adaptive Entropy Coding of Graph Transform Coefficients for Point Cloud Attribute Compression
abstract
A point cloud’s attributes constitutes most of its information content. This is why their efficient compression is of great importance when designing a compression scheme. In this paper, the entropy coding stage of an existing compression method is replaced using Context-based Adaptive Binary Arithmetic Coding (CABAC), which is already largely used for video compression applications. Additionally, the DC transform coefficients are quantized with Differential Pulse Code Modulation (DPCM), yielding better performance at lower bit rates. By using predictive coding and coding techniques adapted to the signal’s characteristics, Bjøntegaard Delta rate savings of up to 31.38% are observable in comparison to previously used entropy coding methods.
Thibaut Meyer, Dominik Mehlem, Christian Rohlfing
VCIP3
2023 Python Wrapper for Context-based Adaptive Binary Arithmetic Coding
abstract
A Python wrapper for Context-based Adaptive Binary Arithmetic Coding (CABAC), extracted from the Test Model (VTM) for Versatile Video Coding (VVC), is presented. Besides providing Python access to CABAC, two extensions are proposed: the probability estimation progress for each context can be traced, and sequences of integer values can be coded without designing a dedicated context model set.
Christian Rohlfing, Thibaut Meyer, Jens Schneider 0001, Jan Voges
VCIP1
2023 GVC: efficient random access compression for gene sequence variations
abstract
BACKGROUND: In recent years, advances in high-throughput sequencing technologies have enabled the use of genomic information in many fields, such as precision medicine, oncology, and food quality control. The amount of genomic data being generated is growing rapidly and is expected to soon surpass the amount of video data. The majority of sequencing experiments, such as genome-wide association studies, have the goal of identifying variations in the gene sequence to better understand phenotypic variations. We present a novel approach for compressing gene sequence variations with random access capability: the Genomic Variant Codec (GVC). We use techniques such as binarization, joint row- and column-wise sorting of blocks of variations, as well as the image compression standard JBIG for efficient entropy coding. RESULTS: Our results show that GVC provides the best trade-off between compression and random access compared to the state of the art: it reduces the genotype information size from 758 GiB down to 890 MiB on the publicly available 1000 Genomes Project (phase 3) data, which is 21% less than the state of the art in random-access capable methods. CONCLUSIONS: By providing the best results in terms of combined random access and compression, GVC facilitates the efficient storage of large collections of gene sequence variations. In particular, the random access capability of GVC enables seamless remote data access and application integration. The software is open source and available at https://github.com/sXperfect/gvc/ .
Yeremia Gunawan Adhisantoso, Jan Voges, Christian Rohlfing, Viktor Tunev, Jens-Rainer Ohm, Jörn Ostermann
BMC Bioinform.3
2022 Deep Hashing with Hash Center Update for Efficient Image Retrieval
abstract
In this paper, we propose an approach for learning binary hash codes for image retrieval. Canonical Correlation Analysis (CCA) is used to design two loss functions for training a neural network such that the correlation between the two views to CCA is maximum. The main motivation for using CCA for feature space learning is that dimensionality reduction is possible and short binary codes could be generated. The first loss maximizes the correlation between the hash centers and the learned hash codes. The second loss maximizes the correlation between the class labels and the classification scores. In this paper, a novel weighted mean and thresholding-based hash center update scheme for adapting the hash centers is proposed. The training loss reaches the theoretical lower bound of the proposed loss functions, showing that the correlation coefficients are maximized during training and substantiating the formation of efficient feature space for retrieval. The measured mean average precision shows that the proposed approach outperforms other state-of-the-art methods.
Abin Jose, Daniel Filbert, Christian Rohlfing, Jens-Rainer Ohm
ICASSP3
2022 Attribute-aware Partitioning for Graph-based Point Cloud Attribute Coding
abstract
The unstructured nature of point cloud data makes compression of their attributes very challenging. In this paper, the known approach of using the Graph Fourier Transform on partitions of the point cloud is improved. It is proposed to make the partitioning process both geometry and attribute-aware, taking all of the point cloud’s characteristics into account simultaneously. Additional information, that allows the decoder to reproduce the partitioning of the encoder, is added to the bitstream. Furthermore, a refinement algorithm which re-estimates the partitioning information at the encoder with the decoder in mind is proposed. Experiments show that the baseline method is outperformed in Bjøntegaard Delta rate reduction by 2.39%, reaching as much as 3.58% at high bitrates.
Thibaut Meyer, Maria Meyer, Dominik Mehlem, Christian Rohlfing
PCS4
2022 Optimized feature space learning for generating efficient binary codes for image retrieval
abstract
In this paper, a novel approach for learning a low-dimensional optimized feature space for image retrieval with minimum intra-class variance and maximum inter-class variance is proposed. The classical approach of Linear Discriminant Analysis (LDA) is generally used for generating an optimized low-dimensional feature space for single-labeled images. Since image retrieval involves images with multiple objects, LDA cannot be directly used for dimensionality reduction and feature space optimization. This problem is addressed by utilizing the relationship between LDA and Canonical Correlation Analysis (CCA) eigenvalues to generate an optimized feature space for both single-labeled and multi-labeled images. A CCA-based network architecture which correlates the low-dimensional feature vectors with the image label vectors is proposed. We design a novel loss function such that the correlation coefficients of CCA are maximized. Our experiments prove that we could train the neural network to reach the theoretical lower bound of loss corresponding to the negative sum of the correlation coefficients. Once the optimized feature space is generated, feature vectors are binarized with the Iterative Quantization (ITQ) approach. Finally, we propose an ensemble network to generate binary codes of desired bit length for retrieval. The measurement of mean average precision shows that the proposed approach outperforms the retrieval results of other single-labeled and multi-labeled image retrieval benchmarks at same bit numbers in a considerable number of cases.
Abin Jose, Erik Stefan Ottlik, Christian Rohlfing, Jens-Rainer Ohm
Signal Process. Image Commun.3
2021 3D Geometry-Based Global Motion Compensation For VVC
abstract
Since most 2D videos are initially captured in 3D environments, developing 3D motion models for video coding is beneficial. This paper introduces a method for extracting 3D geometry data from 2D videos and synthesizing 3D-based virtual Reference Pictures (RPs). These novel RPs are offered to the Versatile Video Coding (VVC) encoder for use in motion compensation. The proposed method generates 3D geometry data in the form of 3D meshes at the encoder and transmits it to the decoder as an overhead bitstream. However, this overhead could eat up the whole coding gain or even make it worse than the anchor VVC. This paper solves this problem by employing 3D mesh processing techniques, e.g., mesh de-noising, mesh decimation, and mesh compression. Simulation results show that the proposed method outperforms VVC up to ~ 3.8%.
Hossein Bakhshi Golestani, Johannes Sauer, Christian Rohlfing, Jens-Rainer Ohm
ICIP3
2021 Sparse Coding-based Intra Prediction in VVC
abstract
Intra prediction is crucial to video coding as it is the only option for prediction when motion compensation either fails or no reference frames are available. Hence, current state-of-the-art video coding standards utilize numerous intra prediction modes in order to predict different angled structures as well as smooth areas. This contribution introduces an additional mode for intra prediction which is based on the concepts of Dictionary Learning (DL), Sparse Coding (SC) [1] and adjusted Anchored Neighborhood Regression (ANR) [2] to be able to adapt to more arbitrary structures. The general idea is built on trained dictionaries, which sparsely represent the reference area of a block to be predicted. Alongside learning the dictionaries, linear projection matrices, projecting the reference areas to the corresponding blocks, are trained with ANR. For the actual intra prediction step, each given reference area is then projected onto the to-be-predicted block by multiple linear projections, which are blended according to the sparse codes representing the reference area. Experimentally, offering the proposed mode to the state-of-the-art video coding standard Versatile Video Coding (VVC) outperforms the traditional VVC modes: In particular, -0.26% BD-rate gains in comparison to the VVC reference software VTM-9.3 and a usage percentage of 12.83% can be achieved on average for the All Intra (AI) coding configuration. Furthermore, a peak coding gain of -0.6% and a usage percentage of 26.66% is observed for the same setup.
Jens Schneider 0001, Dominik Mehlem, Maria Meyer, Christian Rohlfing
PCS4
2021 Dictionary Learning-based Reference Picture Resampling in VVC
abstract
Versatile Video Coding (VVC) introduces the con-cept of Reference Picture Resampling (RPR), which allows for a resolution change of the video during decoding, without introducing an additional Intra Random Access Point (IRAP) into the bitstream. When the resolution is increased, an upsampling operation of the reference picture is required in order to apply motion compensated prediction. Conceptually, the upsampling by linear interpolation filters fails to recover frequencies which were lost during downsampling. Yet, the quality of the upsampled reference picture is crucial to the pre-diction performance. In recent years, machine learning based Super-Resolution (SR) has shown to outperform conventional interpolation filters by far in regard to super-resolving a previ-ously downsampled image. In particular, Dictionary Learning-based Super-Resolution (DLSR) was shown to improve the inter-layer prediction in SHVC [1]. Thus, this paper introduces DLSR to the prediction process in RPR. Further, the approach is experimentally evaluated by an implementation based on the VTM-9.3 reference software. The simulation results show a reduction of the instantaneous bitrate of 0.98% on average at the same objective quality in terms of PSNR. Moreover, the peak bitrate reduction is measured to 4.74% for the “Johnny” sequence of the JVET test set.
Jens Schneider 0001, Christian Rohlfing
VCIP2
2020 Adaptive Resolution Change Using Uncoded Areas and Dictionary Learning-Based Super-Resolution in Versatile Video Coding
abstract
The concept of Adaptive Resolution Change (ARC) in video coding is already known from former international standards such as MPEG-4 [1]. However, in MPEG-4 linear filters are used for upsampling, which is crucial to coding video at varying resolution. With the rise of machine learning-based super-resolution methods in the last decade, powerful algorithms outperforming conventional upsampling methods were developed. This contribution introduces an ARC concept using un-coded areas within a frame of a video sequence and a Dictionary Learning (DL)-based Super-Resolution (SR) scheme. In this concept, a frame contains a picture at different resolution levels, which are spatially separated by the use of slices and tiles. Generally, a tile can be marked as coded or un-coded. Slices which do not hold any coded tiles are omitted from the bitstream. Thus, only one resolution level needs to be coded, while the other is generated at the decoder side. At the encoder a rate-distortion decision is made in order to decide, which resolution level should be coded. Simulation results show that gains with respect to the Versatile Video Coding (VVC) standard in development can be achieved at low bitrates.
Jens Schneider 0001, Johannes Sauer, Christian Rohlfing
ICASSP3
2020 Optimized Convolutional Neural Networks for Video Intra Prediction
abstract
Based on a previously published neural network-based video intra prediction approach, this paper proposes and evaluates several extensions of both the training process as well as the network architectures. In particular, the influence of coding artifacts in the training samples as well as the effect of using different approximations of the residual coding costs as loss functions are investigated. In addition, the architecture is optimized and extended by final deconvolutional layers. Combined with the use of network pruning, it was not only possible to increase the achieved compression gain in comparison to the previous work, but also to decrease the needed number of floating point operations per pixel by more than 72% at the same time.
Maria Meyer, Jonathan Wiesner, Christian Rohlfing
ICIP3
2019 Convolutional Neural Networks for Video Intra Prediction Using Cross-component Adaptation
abstract
Recently, neural networks were shown to improve video and image intra prediction significantly. In this paper, the properties of different architectures for neural network-based intra prediction are evaluated. This includes an analysis of the properties of convolutional neural networks used for this purpose, showing that they outperform fully connected ones especially for complex and low resolution content. Also, the usage of separate networks for luma and chroma prediction, which are able to perform a learned cross-component prediction, is proposed as this is clearly beneficial for the prediction quality. Furthermore, a new way of signaling a neural network-based intra prediction mode in HEVC is investigated. In total this improves the compression performance in terms of average BD-rate changes by -2.0% for the luma and by -1.5% for the chroma channels.
Maria Meyer, Jonathan Wiesner, Jens Schneider 0001, Christian Rohlfing
ICASSP4
2019 Low-Complexity Geometric Inter-Prediction for Versatile Video Coding
abstract
Non-rectangular block partitioning is a well-known method for improved inter-picture prediction in video coding, enabling better spatial adaptation to the signal properties. This contribution presents the most recent proposal of geometric inter-prediction (GIP) made to the Versatile Video Coding (VVC) standardization activity led by the Joint Video Experts Team (JVET). Implemented in the latest test model VTM-5.0 and evaluated according to the JVET Common Test Conditions, the proposed low-complexity GIP scheme provides objective luma BD-rate reductions of 0.22 % for random access and 0.44 % for low-delay test cases at 7% encoder runtime increase and negligible decoder runtime increase. The coding gain is provided by non-triangular partitioned blocks and in the presence of multiple other VVC coding tools. Furthermore, BD-rate reductions of 2.58 % and 2.78 % can be achieved specifically for pure screen content by employing an adaptive blending filter.
Max Bläser, Han Gao 0001, Semih Esenlik, Elena Alshina, Zhijie Zhao, Christian Rohlfing, Eckehard G. Steinbach
PCS6
2019 Reference Picture Synthesis for Video Sequences Captured with a Monocular Moving Camera
abstract
Inter-frame prediction plays an important role in video coding by predicting the current frame from previously encoded pictures, called reference pictures. In the case of camera motion, the content of a current frame could be very different from its reference pictures and may consequently lead to a more difficult Motion Compensation (MC). The main idea of this paper is to process the input 2D video sequence in order to estimate the 3D geometry of the scene and then employ this data to virtually synthesize "geometrically compensated" reference pictures. Since these virtual reference pictures are more similar to the current frame, motion estimation and consequently coding efficiency could be enhanced. The proposed method is tested over six different video sequences and around 11% bitrate reduction is achieved compared to the High Efficiency Video Coding (HEVC) standard.
Hossein Bakhshi Golestani, Christian Rohlfing, Jens-Rainer Ohm
VCIP2
2018 Adaptive Coding of Non-Negative Factorization Parameters with Application to Informed Source Separation
abstract
Informed source separation (ISS) uses source separation for extracting audio objects out of their downmix given some pre-computed parameters. In recent years, non-negative tensor factorization (NTF) has proven to be a good choice for compressing audio objects at an encoding stage. At the decoding stage, these parameters are used to separate the downmix with Wiener-filtering. The quantized NTF parameters have to be encoded to a bit stream prior to transmission. In this paper, we propose to use context-based adaptive binary arithmetic coding (CABAC) for this task. CABAC is widely used in the video coding community and exploits local signal statistics. We adapt CABAC to the task of NTF-based ISS and show that our contribution outperforms reference coding methods.
Max Bläser, Christian Rohlfing, Yingbo Gao, Mathias Wien
ICASSP2
2018 Audio Source Separation with Magnitude Priors: The Beads Model
abstract
Audio source separation comes with the need to devise multichannel filters that can exploit priors about the target signals. In that context, experience shows that modeling magnitude spectra is effective. However, devising a probabilistic model on complex spectral data with a prior on magnitudes is non trivial, because it should both reflect the prior but also be tractable for easy inference. In this paper, we approximate the ideal donut-shaped distribution of a complex variable with approximately known magnitude as a Gaussian mixture model called BEADS (Bayesian Expansion Approximating the Donut Shape) and show that it permits straightforward inference and filtering while effectively constraining the magnitudes of the signals to comply with the prior. As a result, we demonstrate large improvements over the Gaussian baseline for multichannel audio coding when exploiting the BEADS model.
Antoine Liutkus, Christian Rohlfing, Antoine Deleforge
ICASSP2
2017 Very low bitrate spatial audio coding with dimensionality reduction
abstract
In this paper, we show that tensor compression techniques based on randomization and partial observations are very useful for spatial audio object coding. In this application, we aim at transmitting several audio signals called objects from a coder to a decoder. A common strategy is to transmit only the downmix of the objects along some small information permitting reconstruction at the decoder. In practice, this is done by transmitting compressed versions of the objects spectrograms and separating the mix with Wiener filters. Previous research used nonnegative tensor factorizations in this context, with bitrates as low as 1 kbps per object. Building on recent advances on tensor compression, we show that the computation time for encoding can be extremely reduced. Then, we demonstrate how the mixture can be exploited at the decoder to avoid the transmission of many parameters, permitting bitrates as low as 0.1 kbps per object for comparable performance.
Christian Rohlfing, Jérémy E. Cohen, Antoine Liutkus
ICASSP1
2017 Quantization-aware parameter estimation for audio upmixing
abstract
Upmixing consists in extracting audio objects out of their downmix, given some parameters computed beforehand at a coding stage. It is an important task in audio processing with many applications in the entertainment industry. One particularly successful approach for this purpose is to compress the audio objects through nonnegative matrix factorization (NMF) parameters at the coder, to be used for separating the downmix at the decoder. In this paper, we focus on such NMF methods for audio compression, which operate at very low parameter bitrates. In existing methods, parameter estimation and quantization are conducted independently. Here, we propose two extensions: first, we jointly estimate and quantize the parameters at the coder to ensure good reconstruction at the decoder. Second, we propose a parameter refinement method operated at the decoder, that benefits from priors induced by quantization to yield better performance. We show that our contributions outperform existing baseline methods.
Christian Rohlfing, Antoine Liutkus, Julian Mathias Becker
ICASSP1
2016 NMF-based informed source separation
abstract
Informed Source Separation (ISS) is a topic unifying the research fields of both source separation and source coding. Its main objective is to recover audio objects out of a mixture with a source separation step assisted by a set of compact parameters extracted with complete knowledge of the sources. ISS can be used for applications such as active listening and remixing of music (e.g. karaoke). In this paper, we propose a new ISS method which includes a semi-blind source separation (SBSS) step in the ISS decoder to decrease the amount of parameter bit rate. SBSS is conducted by factorizing the mixture in time-frequency domain by nonnegative matrix factorization (NMF). The transmitted parameters consist of a compact NMF initialization as well as residuals calculated in the NMF domain. We show in simulations that using SBSS in the decoder increases the separation quality and that our scheme improves the rate-distortion performance in comparison to a state-of-the art method.
Christian Rohlfing, Julian Mathias Becker, Mathias Wien
ICASSP1
2014 Custom sized non-negative matrix factor deconvolution for sound source separation
abstract
Non-negative Matrix Factorization (NMF) is frequently used for audio source separation. One downside of the NMF is, that it is not able to capture temporal structure of sound events. NMF splits these events into different components. In this paper we present an extension to NMF, which is capable of representing sound events with temporal structure in only one component. We also present an algorithm, which uses this method efficiently. We show that this algorithm leads to a more compact factorization (i.e. with less components) compared to NMF, without losing separation quality.
Julian Mathias Becker, Christian Rohlfing
ICASSP2
2012 Logarithmic Cubic Vector Quantization: Concept and analysis
Christian Rohlfing, Hauke Krüger, Peter Vary
ISITA1