EDBT 2026 Demo / reviewers in the wild / expert
Christian R. Helmrich
dblp:54/9874
· DBLP profile ↗
26ranked-venue papers
12as first author
11since 2021 · last 2025
0000-0001-6606-112XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 12 first-author · 11 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient MDCT-Based Multi-Channel Coding with Perceptual Whitening and Broadband ILD CompensationabstractStereo and multi-channel audio coding for mobile communication applications is challenging due to bitrate, latency and computational complexity constraints. To address these constraints, a novel MDCT-based stereo coding scheme is introduced. A key feature of this scheme is the broadband inter-channel level compensation of the spectrally and temporally whitened channels, followed by a robust mechanism for selecting between Mid/Side and Left/Right coding for each frequency sub-band. For lower bitrates, inter-channel time and phase difference compensation methods are additionally utilized. Enhancements comprising stereo coding of spectral and temporal noise shaping parameters and stereo-aware bandwidth extension further increase coding efficiency. The paper outlines the fundamental concepts of this MDCT-based stereo coding incorporated into the newly introduced 3GPP IVAS codec and their extension to multi-channel coding. Finally, the IVAS codec is compared to state-of-the-art counterparts, demonstrating its advantages in communication scenarios across a wide range of content types. Goran Markovic, Eleni Fotopoulou, Jan Frederik Kiene, Christian R. Helmrich |
ICASSP | 4 |
| 2024 | Fast Constant-Quality Video Encoding Using VVENC With Rate Capping Based On Pre-Analysis StatisticsabstractVVenC, an open Versatile Video Coding (VVC) encoder, has recently been equipped with rate capping functionality in its two-pass rate control modes, providing constrained variable bitrate coding governed by target rate and maximum rate parameters. This paper reports on implementations and evaluation results of straightforward extensions to VVenC which enable the use of the maximum rate parameter also in the single-pass fixed-QP modes, controlled by a base quantization parameter (QP) instead of a target rate. The rate capping in the fixed-QP mode is achieved, with sufficient accuracy, by evaluating only already calculated pre-processing statistics, thereby avoiding increases in encoder runtime. This encoding mode, given that it supports visual quality optimizations such as XPSNR based block-wise perceptual QP adaptation, can be considered a rate capped constant-quality mode, which was missing in VVenC and which is an interesting configuration for video streaming. Christian R. Helmrich, Valeri George, Vignesh V. Menon, Adam Wieckowski, Benjamin Bross, Detlev Marpe |
ICIP | 1 |
| 2024 | Convex-Hull Estimation using Xpsnr for Versatile Video CodingabstractAs adaptive streaming becomes crucial for delivering high-quality video content across diverse network conditions, accurate metrics to assess perceptual quality are essential. This paper explores using the eXtended Peak Signal-to-Noise Ratio (XPSNR) metric as an alternative to the popular Video Multimethod Assessment Fusion (VMAF) metric for determining optimized bitrate-resolution pairs in the context of Versatile Video Coding (VVC). Our study is rooted in the observation that XPSNR shows a superior correlation with subjective quality scores for VVC-coded Ultra-High Definition (UHD) content compared to VMAF. We predict the average XPSNR of VVC-coded bitstreams using spatiotemporal complexity features of the video and the target encoding configuration and then determine the convex-hull online. On average, the proposed convex-hull using XPSNR (VEXUS) achieves an overall quality improvement of 5.84 dB PSNR and 0.62 dB XPSNR while maintaining the same bitrate, compared to the default UHD encoding using the VVenC encoder, accompanied by an encoding time reduction of 44.43% and a decoding time reduction of 65.46%. This shift towards XPSNR as a guiding metric shall enhance the effectiveness of adaptive streaming algorithms, ensuring an optimal balance between bitrate efficiency and perceptual fidelity with advanced video coding standards. Vignesh V. Menon, Christian R. Helmrich, Adam Wieckowski, Benjamin Bross, Detlev Marpe |
ICIP | 2 |
| 2024 | Fast First Pass in Two-Pass Video Encoding Using Sub-SamplingabstractRate control (RC), specifically two-pass, is the main operation mode in VVenC, an open and optimized Versatile Video Coding (VVC) encoder. VVC offers substantial bitrate savings over its predecessor, High Efficiency Video Coding (HEVC), at the price of increased complexity. This complexity increase is apparent in both encoding passes of VVenC. While the complexity redaction in the final pass has been discussed, this paper considers complexity reduction in the first pass, in addition to its already reduced search space. To reduce the overall runtime of a two-pass RC method, spatial and temporal sub-sampling of the first encoding pass is proposed. The experimental results show that the proposed first-pass sub-sampling in two-pass RC can speed up the encoding process of the default two-pass rate control algorithm in VVenC by 18%, with 0.48% loss in coding efficiency, when using the faster preset. Using temporal sub-sampling for the look-ahead, one-pass RC in VVenC can achieve time savings of 11% for bit-rate increases of 0.28%. Anastasia Henkel, Christian R. Helmrich, Tobias Hinz, Jens Brandenburg, Adam Wieckowski, Benjamin Bross, Detlev Marpe, Thomas Wiegand 0001 |
PCS | 2 |
| 2023 | Finalization of VVenC's Screen Content Detector and Two-Pass Rate Control Using Pre-Filtering StatisticsabstractFor improved performance, practical video encoders integrate algorithms for screen content detection and rate control. This paper outlines recently implemented optimizations to both the screen content classifier (SCC) and two-pass rate control (RC) of VVenC, an open Versatile Video Coding (VVC) compliant encoder. The improvements, confirmed by evaluation experiments in random-access configurations using an extended set of test videos, are mainly achieved by leveraging motion error statistics acquired during motion compensated temporal pre-filtering (MCTPF), carried out in VVenC’s pre-analysis stage. All three aspects – pre-analysis stage, SCC, and RC – are revisited herein, and the exploitation of MCTPF data is described. Christian R. Helmrich, Anastasia Henkel, Tobias Hinz, Adam Wieckowski, Benjamin Bross, Detlev Marpe |
ICIP | 1 |
| 2023 | All-Intra Rate Control Using Low Complexity Video Features for Versatile Video CodingabstractVersatile Video Coding (VVC) allows for large compression efficiency gains over its predecessor, High Efficiency Video Coding (HEVC). The added efficiency comes at the cost of increased runtime complexity, especially for encoding. It is thus highly relevant to explore all available runtime reduction options. This paper proposes a novel first pass for two-pass rate control in all-intra configuration, using low-complexity video analysis and a Random Forest (RF)-based machine learning model to derive the data required for driving the second pass. The proposed method is validated using VVenC, an open and optimized VVC encoder. Compared to the default two-pass rate control algorithm in VVenC, the proposed method achieves around 32% reduction in encoding time for the preset faster, while on average only causing 2% BD-rate increase and achieving similar rate control accuracy. Vignesh V. Menon, Anastasia Henkel, Prajit T. Rajendran, Christian R. Helmrich, Adam Wieckowski, Benjamin Bross, Christian Timmerer, Detlev Marpe |
ICIP | 4 |
| 2023 | A Constrained Variable Bit Rate (CVBR) Algorithm for VVenC, an Open VVC Encoder ImplementationabstractRate control (RC) schemes allow audio and video encoders to produce bitstreams according to specific overall bitrate constraints. However, when no rate capping is enforced, the instantaneous bitrate may vary strongly and may exceed the target rate by an order of magnitude, potentially causing playback stutter especially in video streaming scenarios. This paper introduces a rate capping extension for the two RC modes in VVenC, an open Versatile Video Coding (VVC) compliant encoder implementation. After a revisit of VVenC’s two-pass RC approach, the algorithmic details of the rate capping model are described. The paper concludes with an objective evaluation of the performance of the RC extension in a random-access configuration. Christian R. Helmrich, Christian Bartnik, Jens Brandenburg, Adam Wieckowski, Benjamin Bross, Detlev Marpe |
VCIP | 1 |
| 2022 | A Scene Change and Noise Aware Rate Control Method for VVenC, An Open VVC Encoder ImplementationabstractContemporary motion picture content, consisting of scenes with different amounts of visual complexity or camera noise, represents demanding input for video encoders operating in rate control (RC) modes. This paper presents improvements to the 2-pass RC method integrated into VVenC, an open VVC encoder implementation, outlined in previous publications. We specifically introduce three extensions to our RC solution: first, frame type adaptation operating near scene cuts, along with an associated simple detector; second, rate stabilization means to allow for more reliable lookahead based 2-pass RC operation in on-the-fly encoding applications; and third, a low-complexity approach for estimating the instantaneous intensity of camera noise or film grain to avoid large variations in bit consumption when encoding individual frames in the final RC pass. Experimental evaluation confirms that these extensions significantly improve both the objective (BD rate) and subjective (visual) RC performance of VVenC especially on challenging video content. Christian R. Helmrich, Christian Bartnik, Jens Brandenburg, Valeri George, Tobias Hinz, Christian Lehmann, Ivan Zupancic, Adam Wieckowski, Benjamin Bross, Detlev Marpe |
PCS | 1 |
| 2022 | An Optimized Temporal Filter Implementation for Practical ApplicationsabstractVVenC, an open and optimized VVC encoder implementation, employs a temporal filter from the literature as a pre-processing step. The filter effectively reduces camera noise from input video, thereby increasing the encoding gain for lossy encoding, at a price of fairly high complexity, further increased by the necessity of consistent application to many pictures. The filter represents one of the most runtime consuming processing steps for the fastest operating points of VVenC. In this paper, steps are described to reduce the complexity of the temporal filtering in VVenC, to allow its application with low-complexity presets. Overall, the filter runtime is reduced by a factor of around 17 compared to the state of the art, while slightly improving its performance. An additional 4 times speedup is achieved using vectorized implementation. In the proposed version, for the VVenC preset faster, the filter provides 7.36% BD-rate gain at only 2% runtime overhead. Adam Wieckowski, Tobias Hinz, Christian R. Helmrich, Benjamin Bross, Detlev Marpe |
PCS | 3 |
| 2021 | Visually Optimized Two-Pass Rate Control for Video Coding Using the Low-Complexity XPSNR ModelabstractTwo-pass rate control (RC) schemes have proven useful for generating low-bitrate video-on-demand or streaming catalogs. Visually optimized encoding particularly using latest-generation coding standards like Versatile Video Coding (VVC), however, is still a subject of intensive study. This paper describes the two-pass RC method integrated into version 1 of VVenC, an open VVC encoding software. The RC design is based on a novel two-step rate-quantization parameter (R-QP) model to derive the second-pass coding parameters, and it uses the low-complexity XPSNR visual distortion measure to provide numerically as well as visually stable, perceptually R-D optimized encoding results. Random-access evaluation experiments confirm the improved objective as well as subjective performance of our RC solution. Christian R. Helmrich, Ivan Zupancic, Jens Brandenburg, Valeri George, Adam Wieckowski, Benjamin Bross |
VCIP | 1 |
| 2021 | Quantization and Entropy Coding in the Versatile Video Coding (VVC) StandardabstractThe paper provides an overview of the quantization and entropy coding methods in the Versatile Video Coding (VVC) standard. Special focus is laid on techniques that improve coding efficiency relative to the methods included in the High Efficiency Video Coding (HEVC) standard: The inclusion of trellis-coded quantization, the advanced context modeling for entropy coding of transform coefficient levels, the arithmetic coding engine with multi-hypothesis probability estimation, and the joint coding of chroma residuals. Beside a description of the design concepts, the paper also discusses motivations and implementation aspects. The effectiveness of the quantization and entropy coding methods specified in VVC is validated by experimental results. Heiko Schwarz, Muhammed Z. Coban, Marta Karczewicz, Tzu-Der Chuang, Frank Bossen, Alexander Alshin, Jani Lainema, Christian R. Helmrich, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2020 | Xpsnr: A Low-Complexity Extension of The Perceptually Weighted Peak Signal-To-Noise Ratio For High-Resolution Video Quality AssessmentabstractThe objective PSNR metric is known to correlate quite poorly with subjective assessments of video coding quality. Thus, a number of alternative VQA measures such as (MS-)SSIM and VMAF have been proposed. These, however, are often algorithmically complex and difficult to use for visually motivated encoder optimization tasks, especially subjectively optimized bit allocation. In this paper we show that, by way of low-complexity enhancements of our previous work on a perceptually weighted PSNR (WPSNR) metric, addressing shortcomings with video and ultra high-definition content, the prediction of human judgments of video coding quality by the WPSNR can be improved. In fact, the resulting XPSNR seems to match the performance of the aforementioned state-of-the-art methods. Christian R. Helmrich, Mischa Siekmann, Sören Becker 0001, Sebastian Bosse, Detlev Marpe, Thomas Wiegand 0001 |
ICASSP | 1 |
| 2020 | Video Compression Using Generalized Binary Partitioning, Trellis Coded Quantization, Perceptually Optimized Encoding, and Advanced Prediction and Transform CodingabstractIn this paper, we describe a video coding design that enables a higher coding efficiency than the HEVC standard. The proposed video codec follows the design of block-based hybrid video coding, but includes a number of advanced coding tools. A part of the incorporated advanced concepts was developed by the Joint Video Exploration Team, while others are newly proposed. The key aspects of these newly proposed tools are the following. A video frame is subdivided into rectangles of variable size using a binary partitioning with variable split ratios. Three new approaches for generating spatial intra prediction signals are supported: A line-wise application of conventional intra prediction modes, coupled with a mode-dependent processing order, a region-based template matching prediction method and intra prediction modes based on neural networks. For motion-compensated prediction, a multi-hypothesis mode with more than two motion hypotheses can be used. In transform coding, mode dependent combinations of primary and secondary transforms are applied. Moreover, scalar quantization is replaced by trellis-coded quantization and the entropy coding of the quantized transform coefficients is improved. The intra and inter prediction signals can be filtered using an edge-preserving diffusion filter or a non-linear DCT-based thresholding operation. The video codec includes an adaptive in-loop filter for which one of three classifiers can be chosen on a picture basis. We also incorporated an optional encoder control, which adjusts the quantization parameters based on a perceptually motivated distortion measure. In a random access scenario, our proposed video codec achieves luma BD-rate savings between 32.5% for HDR HLG UHD and 39.6% for SDR UHD over the HEVC (HM software) anchor for different categories of test sequences. Jonathan Pfaff, Heiko Schwarz, Detlev Marpe, Benjamin Bross, Santiago De-Luxán-Hernández, Philipp Helle, Christian R. Helmrich, Tobias Hinz, Wang-Q Lim, Jackie Ma, Tung Nguyen 0001, Jennifer Rasch, Michael Schäfer 0003, Mischa Siekmann, Gayathri Venugopal, Adam Wieckowski, Martin Winken, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2019 | Perceptually Optimized Bit-Allocation and Associated Distortion Measure for Block-Based Image or Video CodingabstractIt is well known that input-invariant quantization in perceptual image or video coding often leads to visually suboptimal results and that quantization parameter adaptation (QPA) based on a model of the human visual system can improve subjective coding quality. This paper introduces a simple low-complexity QPA algorithm, controlled using a block-wise perceptually weighted distortion measure representing a generalization of the PSNR metric. The weighting scheme of this WPSNR metric is based on a psychovisual model. It directly leads to a perceptually adapted scaling of the block-wise Lagrange parameter used in the bit-allocation process in the encoder and, consequently, to a block-wise QPA. Unlike prior QPA approaches, the proposal avoids classifications of picture regions and easily extends from still-image or grayscale to video or chromatic coding. The WPSNR metric also uses fewer algorithmic operations than e. g. the multiscale structural similarity measure (MS-SSIM). Due to the results of two formal subjective tests indicating its visual benefit, the QPA proposal has been adopted into VTM, the currently developed Versatile Video Coding (VVC) reference software. Christian R. Helmrich, Sebastian Bosse, Mischa Siekmann, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
DCC | 1 |
| 2019 | Neural Network Guided Perceptually Optimized Bit-Allocation for Block-Based Image and Video CompressionabstractBit-allocation based on the MSE is computationally convenient in image and video compression, but leads to perceptually suboptimal compression results. Distortion sensitivity, modeled as a reference specific property, can be used to improve the accuracy of perceptual quality prediction based on the MSE. This paper shows how distortion sensitivity directly leads to computationally beneficial perceptual optimization of irrelevance reduction and, thereby, of bit-allocation in image and video compression. To this end distortion sensitivity is estimated using a deep convolutional neural network. The proposed method of distortion sensitive bit-allocation is evaluated experimentally using HEVC and on our testset shows average bit-rate reductions with regard to the MOS of 15.9% compared to constant QP-based bit-allocation and 7.3% compared to state-of-the-art perceptual bit-allocation schemes. Sebastian Bosse, Michael Dietzel, Sören Becker 0001, Christian R. Helmrich, Mischa Siekmann, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 4 |
| 2019 | A Study of the Perceptually Weighted Peak Signal-To-Noise Ratio (WPSNR) for Image CompressionabstractThe peak signal-to-noise ratio (PSNR) is the most used objective measure for assessing perceptual image quality when it comes to image and video compression tasks, despite the fact that it exhibits weak performance in reflecting human perception. To address this problem, many image quality assessment (IQA) methods were proposed, e. g. the structural similarity quality measure (SSIM) and its extension, the multi-scale SSIM (MS-SSIM). In this paper we revisit and evaluate a block-based perceptually weighted PSNR (WPSNR) which calculates weighting factors to capture visual sensitivity of local image regions. We further introduce a sample-based version of WPSNR which determines those sensitivity weights with higher spatial accuracy. These methods are computationally inexpensive compared to other similarity measures and are shown to outperform PSNR, SSIM and similar perceptual quality measures when it comes to approximate subjective ratings of JPEG or JPEG2000 compressed images. Johannes Erfurt, Christian R. Helmrich, Sebastian Bosse, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 2 |
| 2019 | Inter-Component Transform for Color Video CodingabstractIn natural digital images and videos, correlations between color components can be observed. These correlations can be exploited to achieve additional coding gain in modern block-based hybrid video coding. To this end, we propose the use of a block-wise, rotational inter-component transform (ICT) applied to the two residual chroma signals that result from conventional intra or inter-picture prediction. Different ICT parameterizations in terms of number and quantization of the rotational angles as well as resulting components signaled in the coded bitstream are investigated. An implementation into the currently developed Versatile Video Coding (VVC) reference software provides average bitrate savings of up to 0.7% (All Intra configuration) with negligible increases in implementation complexity and runtime. Our proposal has been adopted into the VVC draft specification text. Christian Rudat, Christian R. Helmrich, Jani Lainema, Tung Nguyen 0001, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
PCS | 2 |
| 2016 | Signal-adaptive switching of overlap ratio in audio transform codingabstractContemporary perceptual audio coders, all of which apply the modified discrete cosine transform (MDCT), with an overlap ratio of 50%, for frequency-domain quantization, provide good coding quality even at low bit-rates. However, relatively long frames are required for acceptable low-rate performance also for quasi-stationary harmonic input, leading to increased algorithmic latency and reduced temporal coding resolution. This paper investigates the alternative approach of employing the extended lapped transform (ELT), with 75% overlap ratio, on such input. To maintain a high time resolution for coding of transient segments, the ELT definition is modified such that frame-wise switching between ELT (for quasi-stationary) and MDCT coding (for non-stationary or non-tonal regions), with complete time-domain aliasing cancelation and no increase in frame length, becomes possible. A new ELT window function with improved side-lobe rejection to avoid framing artifacts is also derived. Blind subjective evaluation of the switched-ratio proposal confirms the benefit of the signal-adaptive design. Christian R. Helmrich, Bernd Edler |
ICASSP | 1 |
| 2016 | Audio Coding Using Overlap and Kernel AdaptationabstractPerceptual audio coding schemes typically apply the modified discrete cosine transform (MDCT) with different lengths and windows, and utilize signal-adaptive switching between these on a perframe basis for best subjective performance. In previous papers, the authors demonstrated that further quality gains can be achieved for some input signals using additional transform kernels such as the modified discrete sine transform (MDST) or greater inter-transform overlap by means of a modified extended lapped transform (MELT). This work discusses the algorithmic procedures and codec modifications necessary to combine all of the above features-transform length, window shape, transform kernel, and overlap ratio switching-into a flexible input-adaptive coding system. It is shown that, due to full time-domain aliasing cancelation, this system supports perfect signal reconstruction in the absence of quantization and, thanks to fast realizations of all transforms, increases the codec complexity only negligibly. The results of a 5.1 multichannel listening test are also reported. Christian R. Helmrich, Bernd Edler |
IEEE Signal Process. Lett. | 1 |
| 2015 | Arithmetic coding of speech and audio spectra using tcx based on linear predictive spectral envelopesabstractUnified speech and audio codecs often use a frequency domain coding technique of the transform coded excitation (TCX) type. It is based on modeling the speech source with a linear predictor, spectral weighting by a perceptual model and entropy coding of the frequency components. While previous approaches have used neighbouring frequency components to form a probability model for the entropy coder of spectral components, we propose to use the magnitude of the linear predictor to estimate the variance of spectral components. Since the linear predictor is transmitted in any case, this method does not require any additional side info. Subjective measurements show that the proposed methods give a statistically significant improvement in perceptual quality when the bit-rate is held constant. Consequently, the proposed method has been adopted to the 3GPP Enhanced Voice Services speech coding standard. Tom Bäckström, Christian R. Helmrich |
ICASSP | 2 |
| 2015 | Low delay LPC and MDCT-based audio coding in the EVS codecabstractSpeech coders operating in time domain can be extended with a frequency domain mode to improve encoding of music, even though this is challenging at low delay. In such a scenario, the short analysis window limits the benefit of the transform coder, while a delayless switch between the two coders constrains the system further. The paper presents an LPC and MDCT-based audio coder part of the new 3GPP codec for Enhanced Voice Services, which aims to solve the issues. Several advanced coding tools are introduced to alleviate the constraints: transient handling is improved, harmonic structures are better preserved, and the modeling of the zero-quantized frequencies is enhanced. Test results show that the obtained low-delay switched coder brings a clear improvement over a speech coder and is competitive even in comparison to audio coders with higher delay. Guillaume Fuchs, Christian R. Helmrich, Goran Markovic, Matthias Neusinger, Emmanuel Ravelli, Takehiro Moriya |
ICASSP | 2 |
| 2015 | Spectral envelope reconstruction via IGF for audio transform codingabstractIn low-bitrate audio coding, modern coders often rely on efficient parametric techniques to enhance the performance of the waveform preserving transform coder core. While the latter features well-known perceptually adapted quantization of spectral coefficients, parametric techniques reconstruct the signal parts that have been quantized to zero by the encoder to meet the low-bitrate constraint. Large numbers of zeroed spectral values and especially consecutive zeros constituting gaps often lead to audible artifacts at the decoder. To avoid such artifacts the new 3GPP Enhanced Voice Services (EVS) coding standard utilizes noise filling and intelligent gap filling (IGF) techniques, guided by spectral envelope information. In this paper the underlying considerations of the parametric energy adjustment and transmission in EVS and its relation to noise filling, IGF, and tonality preservation are presented. It is further shown that complex-valued IGF envelope calculation in the encoder improves the temporal energy stability of some signals while retaining real-valued decoder-side processing. Christian R. Helmrich, Andreas Niedermeier, Sascha Disch, Florin Ghido |
ICASSP | 1 |
| 2015 | Low-complexity and robust coding mode decision in the EVS coderabstractSeveral state-of-the-art switched audio codecs employ the closed-loop mode decision to select the best coding mode at every frame. The closed-loop mode selection is known to have good performance but also high complexity. The new approach we propose in this paper is a low-complexity version of the closed-loop approach, based on similar decisions which compute the coding distortion of each mode and select the one with the lowest distortion. Our approach differs mainly in the way the coding distortions are calculated. We are able to notably reduce the complexity by only estimating the distortions without encoding and decoding the input for each mode. The new approach was implemented in the EVS codec standard and evaluated both objectively and subjectively. Compared to the closed-loop approach, it yields similar performance and lower complexity. Emmanuel Ravelli, Christian R. Helmrich, Guillaume Fuchs, Markus Multrus |
ICASSP | 2 |
| 2014 | Improved low-delay MDCT-based coding of both stationary and transient audio signalsabstractGeneral-purpose MDCT-based audio coders like MP3 or HE-AAC utilize long inter-transform overlap and lookahead-based transform length switching to provide good coding quality for both stationary and non-stationary, i. e. transient, input signals even at low bitrates. In low-delay communication scenarios such as Voice over IP, however, algorithmic delay due to framing and overlap typically needs to be reduced and additional lookahead must be avoided. We show that these restrictions limit the performance of contemporary low-delay transform coders on either stationary or transient material and propose 3 modifications: an improved noise substitution technique and increased overlap between “long”transforms for stationary, and “long to short” transform length switching without lookahead and directly from the long overlap for transient frames. A listening test indicates the merit of these changes when integrated into AAC-LD. Christian R. Helmrich, Goran Markovic, Bernd Edler |
ICASSP | 1 |
| 2014 | Decorrelated innovative codebooks for ACELP using factorization of autocorrelation matrix
Tom Bäckström, Christian R. Helmrich |
INTERSPEECH | 2 |
| 2011 | Efficient transform coding of two-channel audio signals by means of complex-valued stereo predictionabstractTraditional MDCT-based perceptual audio coding schemes employ mid/side and intensity stereo techniques to allow efficient joint coding of the two channels of a stereophonic signal. These techniques, however, provide only little coding gain for critical stereo signals characterized by spectral components with a distinct level or phase difference between the channels. To overcome this deficiency, we propose an extension to the mid/side coding paradigm that utilizes complex-valued inter-channel linear prediction in the MDCT spectral domain. The required imaginary spectrum (MDST) is calculated in a computationally efficient manner without additional algorithmic delay. A formal listening test conducted in the course of the ISO/MPEG standardization of the unified speech and audio codec USAC illustrates that the proposed stereo prediction approach pro vides significant improvements in coding efficiency and shows that at 96 kb/s, excellent quality can be obtained even for critical signals. Christian R. Helmrich, Pontus Carlsson, Sascha Disch, Bernd Edler, Johannes Hilpert, Matthias Neusinger, Heiko Purnhagen, Nikolaus Rettelbach, Julien Robilliard, Lars F. Villemoes |
ICASSP | 1 |