EDBT 2026 Demo / reviewers in the wild / expert
Bernd Edler
dblp:76/2413
· DBLP profile ↗
37ranked-venue papers
1as first author
8since 2021 · last 2025
0009-0009-3172-6369ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 since 2021Computer networks · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FlowMAC: Conditional Flow Matching for Audio Coding at Low Bit RatesabstractThis paper introduces FlowMAC, a novel neural audio codec for high-quality general audio compression at low bit rates based on conditional flow matching (CFM). FlowMAC jointly learns a mel spectrogram encoder, quantizer and decoder. At inference time the decoder integrates a continuous normalizing flow via an ODE solver to generate a high-quality mel spectrogram. This is the first time that a CFM-based approach is applied to general audio coding, enabling a scalable, simple and memory efficient training. Our subjective evaluations show that FlowMAC at 3 kbps achieves similar quality as state-of-the-art GAN-based and DDPM-based neural audio codecs at double the bit rate. Moreover, FlowMAC offers a tunable inference pipeline, which permits to trade off complexity and quality. This enables real-time coding on CPU, while maintaining high perceptual quality. Nicola Pia, Martin Strauss 0003, Markus Multrus, Bernd Edler |
ICASSP | 4 |
| 2025 | Frequency Domain Prediction of Tonal Signals With Time-Varying PitchesabstractIn this letter, we propose an Extended Frequency Domain Joint Harmonics Prediction (EFDJHP) algorithm, which does long term prediction directly in the transform domain on tonal signals with time-varying pitches for transform speech and audio coding. EFDJHP is an algorithm extension of a previously proposed Frequency Domain Joint Harmonics Prediction (FDJHP) algorithm where a constant pitch between neighboring frames was assumed (Guo and Edler, 2021). A linearly changing pitch between adjacent frames is assumed in EFDJHP, where the linearity can be assumed to be local and updated across frames. EFDJHP works in a backward prediction fashion without additional algorithmic delay, and needs very few side information for the forward adaption of the predictor coefficients. Bitrate saving analysis and a listening test show that EFDJHP can improve the coding efficiency on tonal signals with frequent pitch variations such as singing voices and speech. Bernd Edler |
IEEE Signal Process. Lett. | 2 |
| 2023 | Predicting Preferred Dialogue-to-Background Loudness Difference in Dialogue-Separated AudioabstractDialogue Enhancement (DE) enables the rebalancing of dialogue and background sounds to fit personal preferences and needs in the context of broadcast audio. When individual audio stems are unavailable from production, Dialogue Separation (DS) can be applied to the final audio mixture to obtain esti-mates of these stems. This work focuses on Preferred Loudness Differences (PLDs) between dialogue and background sounds. While previous studies determined the PLD through a listening test employing original stems from production, stems estimated by DS are used in the present study. In addition, a larger variety of signal classes is considered. PLDs vary substantially across individuals (average interquartile range: 5.7 LU). Despite this variability, PLDs are found to be highly dependent on the signal type under consideration, and it is shown that median PLDs can be predicted using objective intelligibility metrics. Two existing baseline prediction methods - intended for use with original stems - displayed a Mean Absolute Error (MAE) of 7.5 LU and 5 LU, respectively. A modified baseline (MAE: 3.2 LU) and an alternative approach (MAE: 2.5 LU) are proposed. Results support the viability of processing final broadcast mixtures with DS and offering an alternative remixing that accounts for median PLDs. Luca Resti, Martin Strauss 0003, Matteo Torcoli, Emanuël A. P. Habets, Bernd Edler |
QoMEX | 5 |
| 2022 | A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT DomainabstractFrequency domain processing, and in particular the use of Modified Discrete Cosine Transform (MDCT), is the most widespread approach to audio coding. However, at low bitrates, audio quality, especially for speech, degrades drastically due to the lack of available bits to directly code the transform coefficients. Traditionally, post-filtering has been used to mitigate artefacts in the coded speech by exploiting a-priori information of the source and extra transmitted parameters. Recently, datadriven post-filters have shown better results, but at the cost of significant additional complexity and delay. In this work, we propose a mask-based post-filter operating directly in MDCT domain of the codec, inducing no extra delay. The real-valued mask is applied to the quantized MDCT coefficients and is estimated from a relatively lightweight convolutional encoder-decoder network. Our solution is tested on the recently standardized low-delay, low-complexity codec (LC3) at lowest possible bitrate of 16 kbps. Objective and subjective assessments clearly show the advantage of this approach over the conventional post-filter, with an average improvement of 10 MUSHRA points over the LC3 coded speech. Kishan Gupta, Srikanth Korse, Bernd Edler, Guillaume Fuchs |
ICASSP | 3 |
| 2022 | Improved Normalizing Flow-Based Speech Enhancement Using an all-Pole Gammatone Filterbank for Conditional Input RepresentationabstractDeep generative models for Speech Enhancement (SE) received increasing attention in recent years. The most prominent example are Generative Adversarial Networks (GANs), while normalizing flows (NF) received less attention despite their potential. Building on previous work, architectural modifications are proposed, along with an investigation of different conditional input representations. Despite being a common choice in related works, Mel-spectrograms demonstrate to be inadequate for the given scenario. Alternatively, a novel All-Pole Gammatone filterbank (APG) with high temporal resolution is proposed. Although computational evaluation metric results would suggest that state-of-the-art GAN-based methods perform best, a perceptual evaluation via a listening test indicates that the presented NF approach (based on time domain and APG) performs best, especially at lower SNRs. On average, APG outputs are rated as having good quality, which is unmatched by the other methods, including GAN. Martin Strauss 0003, Matteo Torcoli, Bernd Edler |
SLT | 3 |
| 2021 | A Flow-Based Neural Network for Time Domain Speech EnhancementabstractSpeech enhancement involves the distinction of a target speech signal from an intrusive background. Although generative approaches using Variational Autoencoders or Generative Adversarial Networks (GANs) have increasingly been used in recent years, normalizing flow (NF) based systems are still scarse, despite their success in related fields. Thus, in this paper we propose a NF framework to directly model the enhancement process by density estimation of clean speech utterances conditioned on their noisy counterpart. The WaveGlow model from speech synthesis is adapted to enable direct enhancement of noisy utterances in time domain. In addition, we demonstrate that nonlinear input companding benefits the model performance by equalizing the distribution of input samples. Experimental evaluation on a publicly available dataset shows comparable results to current state-of-the-art GAN-based approaches, while surpassing the chosen baselines using objective evaluation metrics. Martin Strauss 0003, Bernd Edler |
ICASSP | 2 |
| 2021 | A Hands-On Comparison of DNNs for Dialog Separation Using Transfer Learning from Music Source SeparationabstractThis paper describes a hands-on comparison on using state-of-the-art music source separation deep neural networks (DNNs) before and after task-specific fine-tuning for separating speech content from non-speech content in broadcast audio (i.e., dialog separation). The music separation models are selected as they share the number of channels (2) and sampling rate (44.1 kHz or higher) with the considered broadcast content, and vocals separation in music is considered as a parallel for dialog separation in the target application domain. These similarities are assumed to enable transfer learning between the tasks. Three models pre-trained on music (Open-Unmix, Spleeter, and Conv-TasNet) are considered in the experiments, and fine-tuned with real broadcast data. The performance of the models is evaluated before and after fine-tuning with computational evaluation metrics (SI-SIRi, SI-SDRi, 2f-model), as well as with a listening test simulating an application where the non-speech signal is partially attenuated, e.g., for better speech intelligibility. The evaluations include two reference systems specifically developed for dialog separation. The results indicate that pre-trained music source separation models can be used for dialog separation to some degree, and that they benefit from the fine-tuning, reaching a performance close to task-specific solutions. Martin Strauss 0003, Jouni Paulus, Matteo Torcoli, Bernd Edler |
Interspeech | 4 |
| 2021 | Frequency Domain Long-Term Prediction for Low Delay General Audio CodingabstractIn this paper we propose a long-term prediction method for low delay transform domain general audio coders. This Frequency Domain Joint Harmonics Prediction (FDJHP) method operates directly in the Modified Discrete Cosine Transform (MDCT) domain and can enhance the coding efficiency, even under very low frequency resolutions. We compare this new method with state-of-the-art MDCT based methods by analyzing bitrate savings and by a listening test using test signals with strong harmonic components. The results indicate that it outperforms an existing method, which also directly operates in the frequency domain. Additionally, we show how it can be combined with the existing techniques into an adaptive system, where the different methods can complement each other. Bernd Edler |
IEEE Signal Process. Lett. | 2 |
| 2019 | Perceptual Audio Coding with Adaptive Non-uniform Time/frequency Tilings Using Subband Merging and Time Domain Aliasing ReductionabstractIn this paper, we investigate the coding efficiency of perceptual coding using an adaptive non-uniform orthogonal filter-bank based on MDCT analysis/synthesis and time domain aliasing reduction. We compare its performance to a system using a traditional adaptive uniform MDCT filterbank with window switching. The comparison is performed using a listening test at two different quantization settings. The statistical evaluation shows that the percetpual quality of the nonuniform filterbank significantly out-performs that of the uniform filterbank by 5 to 10 MUSHRA points. Nils Werner, Bernd Edler |
ICASSP | 2 |
| 2019 | Time-Varying Time-Frequency Tilings Using Non-Uniform Orthogonal Filterbanks Based on MDCT Analysis/Synthesis and Time Domain Aliasing ReductionabstractTime Domain Aliasing Reduction (TDAR) is a method to improve the impulse response compactness of non-uniform orthogonal Modified Discrete Cosine Transforms (MDCT). Previously, TDAR was only possible between frames of identical time-frequency tilings, however in this letter we describe a method to overcome this limitation. This method enables the use of TDAR between two consecutive frames of different time-frequency tilings by introducing another subband merging or subband splitting step. Consecutively, this method allows more flexible and adaptive filterbank tilings while retaining compact impulse responses, two attributes needed for efficient perceptual audio coding. Nils Werner, Bernd Edler |
IEEE Signal Process. Lett. | 2 |
| 2019 | CountNet: Estimating the Number of Concurrent Speakers Using Supervised LearningabstractEstimating the maximum number of concurrent speakers from single-channel mixtures is a challenging problem and an essential first step to address various audio-based tasks such as blind source separation, speaker diarization, and audio surveillance. We propose a unifying probabilistic paradigm, where deep neural network architectures are used to infer output posterior distributions. These probabilities are in turn processed to yield discrete point estimates. Designing such architectures often involves two important and complementary aspects that we investigate and discuss. First, we study how recent advances in deep architectures may be exploited for the task of speaker count estimation. In particular, we show that convolutional recurrent neural networks outperform recurrent networks used in a previous study when adequate input features are used. Even for short segments of speech mixtures, we can estimate up to five speakers, with a significantly lower error than other methods. Second, through comprehensive evaluation, we compare the best-performing method to several baselines, as well as the influence of gain variations, different data sets, and reverberation. The output of our proposed method is compared to human performance. Finally, we give insights into the strategy used by our proposed method. Fabian-Robert Stöter, Soumitro Chakrabarty, Bernd Edler, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Blind Bandwidth Extension Based on Convolutional and Recurrent Deep Neural NetworksabstractA blind bandwidth extension (BBWE) expands the bandwidth of telephone speech which often is limited to 0.2 to 3.4 kHz. The advantage is an increased perceived quality as well as an increased intelligibility. This work presents a BBWE similar to state-of-the-art bandwidth extensions like Intelligent Gap Filling with the difference that all processing is done in the decoder without the need of transmitting extra bits. Parameters like spectral envelope are estimated by a regressive Convolutional Deep Neuronal Network (CNN) with long short-term memory (LSTM). The system operates on frames of 20 ms without additional algorithmic delay and can be applied in state-of-the-art speech and audio codecs. Konstantin Schmidt, Bernd Edler |
ICASSP | 2 |
| 2018 | Classification vs. Regression in Supervised Learning for Single Channel Speaker Count EstimationabstractThe task of estimating the maximum number of concurrent speakers from single channel mixtures is important for various audio-based applications, such as blind source separation, speaker diarisation, audio surveillance or auditory scene classification. Building upon powerful machine learning methodology, we develop a Deep Neural Network (DNN) that estimates a speaker count. While DNNs efficiently map input representations to output targets, it remains unclear how to best handle the network output to infer integer source count estimates, as a discrete count estimate can either be tackled as a regression or a classification problem. In this paper, we investigate this important design decision and also address complementary parameter choices such as the input representation. We evaluate a state-of-the-art DNN audio model based on a Bi-directional Long Short-Term Memory network architecture for speaker count estimations. Through experimental evaluations aimed at identifying the best overall strategy for the task and show results for five seconds speech segments in mixtures of up to ten speakers. Fabian-Robert Stöter, Soumitro Chakrabarty, Bernd Edler, Emanuël A. P. Habets |
ICASSP | 3 |
| 2018 | Single-Channel Dereverberation Using Direct MMSE Optimization and Bidirectional LSTM NetworksabstractDereverberation is useful in hands-free communication and voice controlled devices for distant speech acquisition. Single-channel dereverberation can be achieved by applying a time-frequency (TF) mask to the short-time Fourier transform (STFT) representation of a reverberant signal. Recent approaches have used deep neural networks (DNNs) to estimate such masks. Previously proposed DNN-based mask estimation methods train a DNN to minimize the mean-squared-error (MSE) between the desired and estimated masks. Recent TF mask estimation methods for signal separation directly minimize instead the MSE between the desired and estimated STFT magnitudes. We apply this direct optimization concept to dereverberation. Moreover, as reverberation exceeds the duration of a single STFT frame, we propose to use a bidirectional long short-term memory (LSTM) network which is able to take the relation between multiple STFT frames into account. We evaluated our method for different reverberation times and source-microphone distances using simulated as well as measured room impulse responses of different rooms. An evaluation of the proposed method and a comparison with a state-of-the-art method demonstrate the superiority of our approach and its robustness to different acoustic conditions. Wolfgang Mack, Soumitro Chakrabarty, Fabian-Robert Stöter, Sebastian Braun, Bernd Edler, Emanuël A. P. Habets |
INTERSPEECH | 5 |
| 2017 | Nonuniform Orthogonal Filterbanks Based on MDCT Analysis/Synthesis and Time-Domain Aliasing ReductionabstractIn this letter we describe nonuniform orthogonal modified discrete cosine transform (MDCT) filterbanks and time-domain aliasing reduction (TDAR). By adding a postprocessing step to the MDCT, our method allows for arbitrary nonuniform frequency resolutions using subband merging with smooth windowing and overlap in frequency. This overlap allows for an improved temporal compactness of the impulse response, which is especially useful for audio coders. The postprocessing step comprises another lapped MDCT transform along the frequency axis and TDAR along each subband signal. Nils Werner, Bernd Edler |
IEEE Signal Process. Lett. | 2 |
| 2016 | Signal-adaptive switching of overlap ratio in audio transform codingabstractContemporary perceptual audio coders, all of which apply the modified discrete cosine transform (MDCT), with an overlap ratio of 50%, for frequency-domain quantization, provide good coding quality even at low bit-rates. However, relatively long frames are required for acceptable low-rate performance also for quasi-stationary harmonic input, leading to increased algorithmic latency and reduced temporal coding resolution. This paper investigates the alternative approach of employing the extended lapped transform (ELT), with 75% overlap ratio, on such input. To maintain a high time resolution for coding of transient segments, the ELT definition is modified such that frame-wise switching between ELT (for quasi-stationary) and MDCT coding (for non-stationary or non-tonal regions), with complete time-domain aliasing cancelation and no increase in frame length, becomes possible. A new ELT window function with improved side-lobe rejection to avoid framing artifacts is also derived. Blind subjective evaluation of the switched-ratio proposal confirms the benefit of the signal-adaptive design. Christian R. Helmrich, Bernd Edler |
ICASSP | 2 |
| 2016 | Common fate model for unison source separationabstractIn this paper we present a novel source separation method aiming to overcome the difficulty of modelling non-stationary signals. The method can be applied to mixtures of musical instruments with frequency and/or amplitude modulation, e.g. typically caused by vibrato. It is based on a signal representation that divides the complex spectrogram into a grid of patches of arbitrary size. These complex patches are then processed by a two-dimensional discrete Fourier transform, forming a tensor representation which reveals spectral and temporal modulation textures. Our representation can be seen as an alternative to modulation transforms computed on magnitude spectrograms. An adapted factorization model allows to decompose different time-varying harmonic sources based on their particular common modulation profile: hence the name Common Fate Model. The method is evaluated on musical instrument mixtures playing the same fundamental frequency (unison), showing improvement over other state-of-the-art methods. Fabian-Robert Stöter, Antoine Liutkus, Roland Badeau, Bernd Edler, Paul Magron |
ICASSP | 4 |
| 2016 | Audio Coding Using Overlap and Kernel AdaptationabstractPerceptual audio coding schemes typically apply the modified discrete cosine transform (MDCT) with different lengths and windows, and utilize signal-adaptive switching between these on a perframe basis for best subjective performance. In previous papers, the authors demonstrated that further quality gains can be achieved for some input signals using additional transform kernels such as the modified discrete sine transform (MDST) or greater inter-transform overlap by means of a modified extended lapped transform (MELT). This work discusses the algorithmic procedures and codec modifications necessary to combine all of the above features-transform length, window shape, transform kernel, and overlap ratio switching-into a flexible input-adaptive coding system. It is shown that, due to full time-domain aliasing cancelation, this system supports perfect signal reconstruction in the absence of quantization and, thanks to fast realizations of all transforms, increases the codec complexity only negligibly. The results of a 5.1 multichannel listening test are also reported. Christian R. Helmrich, Bernd Edler |
IEEE Signal Process. Lett. | 2 |
| 2015 | Multi-Sensor Cello Recordings for Instantaneous Frequency EstimationabstractEstimating the fundamental frequency (F0) of a signal is a well studied task in audio signal processing with many applications. If the F0 varies over time, the complexity increases, and it is also more difficult to provide ground truth data for evaluation. In this paper we present a novel dataset of cello recordings addressing the lack of reference annotations for musical instruments. Besides audio data, we include sensor recordings capturing the finger position on the fingerboard which is converted into an instantaneous frequency estimate. In speech processing, the electroglottograph (EGG) is able to capture the excitation signal of the vocal tract, which is then used to generate a reference instantaneous F0. Inspired by this approach, we included high speed video camera recordings to extract the excitation signal originating from the moving string. The derived data can be used to analyze vibratos --- a very commonly used playing style. The dataset is released under a Creative Commons license. Fabian-Robert Stöter, Michael G. Müller, Bernd Edler |
ACM Multimedia | 3 |
| 2014 | Unison Source Separation
Fabian-Robert Stöter, Stefan Bayer, Bernd Edler |
DAFx | 3 |
| 2014 | Improved low-delay MDCT-based coding of both stationary and transient audio signalsabstractGeneral-purpose MDCT-based audio coders like MP3 or HE-AAC utilize long inter-transform overlap and lookahead-based transform length switching to provide good coding quality for both stationary and non-stationary, i. e. transient, input signals even at low bitrates. In low-delay communication scenarios such as Voice over IP, however, algorithmic delay due to framing and overlap typically needs to be reduced and additional lookahead must be avoided. We show that these restrictions limit the performance of contemporary low-delay transform coders on either stationary or transient material and propose 3 modifications: an improved noise substitution technique and increased overlap between “long”transforms for stationary, and “long to short” transform length switching without lookahead and directly from the long overlap for transient frames. A listening test indicates the merit of these changes when integrated into AAC-LD. Christian R. Helmrich, Goran Markovic, Bernd Edler |
ICASSP | 3 |
| 2013 | Improved arithmetic coding for Time-Warped MDCT based audio codingabstractThe Time-Warped Modified Discrete Cosine Transform (TW-MDCT) improves the energy compaction for harmonic signals with varying fundamental frequency compared to the plain MDCT. Adaptive context based entropy coding has the potential to provide higher gain over memoryless entropy coding. But in combination with the TW-MDCT, the context based adaptive coding may lead to suboptimal coding. This paper presents an algorithm for improving the context for the TW-MDCT. This is mainly achieved by exploiting already available information on the frequency variation needed by the TW-MDCT. This results in an improved entropy coding. Stefan Bayer, Bernd Edler |
ICASSP | 2 |
| 2013 | Cheap beeps - Efficient synthesis of sinusoids and sweeps in the MDCT domainabstractModern transform audio coders often employ parametric enhancements, like noise substitution or bandwidth extension. In addition to these well-known parametric tools, it might also be desirable to synthesize parametric sinusoidal tones in the decoder. Low computational complexity is an important criterion in codec development and essential for acceptance and deployment. Therefore, efficient ways of generating these tones are needed. Since contemporary codecs like AAC or USAC are based on an MDCT domain representation of audio, we propose to generate synthetic tones by patching tone patterns into the MDCT spectrum at the decoder. We demonstrate how appropriate spectral patterns can be derived and adapted to their target location in (and between) the MDCT time/frequency (t/f) grid to seamlessly synthesize high quality sinusoidal tones including sweeps. Sascha Disch, Benjamin Schubert, Bernd Edler |
ICASSP | 3 |
| 2011 | Frequency selective pitch transposition of audio signalsabstractModern music production often uses pre-recorded pieces of audio, so-called samples, taken from a huge sample database. Consequently, there is an increasing demand to extensively adapt these samples to their intended new musical environment in a flexible way. Such an application, for instance, retroactively changes the key mode of audio recordings, e.g. from a major key to minor key by a frequency selective transposition of pitch. Recently, the modulation vocoder (MODVOC) has been proposed to handle this task. In this paper, two enhancements to the MODVOC are presented and the subjective quality of its application to selective pitch transposition is assessed. Moreover, the proposed scheme is compared with results obtained by applying a commercial computer program, which became newly available on the market. The proposed method is clearly preferred in terms of the perceptual quality aspect "melody and chords transposition", while the commercial program is favored by the majority with regard to the aspect "timbre preservation". Sascha Disch, Bernd Edler |
ICASSP | 2 |
| 2011 | Efficient transform coding of two-channel audio signals by means of complex-valued stereo predictionabstractTraditional MDCT-based perceptual audio coding schemes employ mid/side and intensity stereo techniques to allow efficient joint coding of the two channels of a stereophonic signal. These techniques, however, provide only little coding gain for critical stereo signals characterized by spectral components with a distinct level or phase difference between the channels. To overcome this deficiency, we propose an extension to the mid/side coding paradigm that utilizes complex-valued inter-channel linear prediction in the MDCT spectral domain. The required imaginary spectrum (MDST) is calculated in a computationally efficient manner without additional algorithmic delay. A formal listening test conducted in the course of the ISO/MPEG standardization of the unified speech and audio codec USAC illustrates that the proposed stereo prediction approach pro vides significant improvements in coding efficiency and shows that at 96 kb/s, excellent quality can be obtained even for critical signals. Christian R. Helmrich, Pontus Carlsson, Sascha Disch, Bernd Edler, Johannes Hilpert, Matthias Neusinger, Heiko Purnhagen, Nikolaus Rettelbach, Julien Robilliard, Lars F. Villemoes |
ICASSP | 4 |
| 2011 | Prediction of DCT coefficients considering motion compensation error distributionsabstractCurrent video coding techniques use a Discrete Cosine Transform (DCT) to reduce spatial correlations within the motion estimation residual. Often the correlation cannot be completely eliminated leaving the transform coefficients statistically dependent. The presented paper proposes a method to predict these coefficients on a block level by using the distribution of the prediction error variance to improve coding efficiency. First experiments lead to a reduction in bit rate by 1.83% when compared to the standard JM 17.2 implementation results. Julia Schmidt, Bernd Edler, Jörn Ostermann |
VCIP | 2 |
| 2009 | Multiband perceptual modulation analysis, processing and synthesis of audio signalsabstractThe decomposition of audio signals into perceptually meaningful multiband modulation components opens up new possibilities for advanced signal processing. The signal adaptive analysis approach proposed in this paper will be shown to provide a powerful handle on the signal's perceptual properties: pitch, timbre or roughness can be manipulated straight forward. Additionally a synthesis method is specified providing high subjective perceptual quality. Furthermore, as an application example, a novel audio processing technique is proposed which changes the key mode of a given piece of music e.g. from major to minor key or vice versa. Sascha Disch, Bernd Edler |
ICASSP | 2 |
| 2007 | Automatic speech recognition with a cochlear implant front-endabstractToday, cochlear implants (CIs) are the treatment of choice in patients with profound hearing loss. However speech intelligibility with these devices is still limited. A factor that determines hearing performance is the processing method used in CIs. Therefore, research is focused on designing different speech processing methods. The evaluation of these strategies is subject to variability as it is usually performed with cochlear implant recipients. Hence, an objective method for the evaluation would give more robustness compared to the tests performed with CI patients. This paper proposes a method to evaluate signal processing strategies for CIs based on a hidden markov model speech recognizer. Two signal processing strategies for CIs, the Advanced Combinational Encoder (ACE) and the Psychoacoustic Advanced Combinational Encoder (PACE), have been compared in a phoneme recognition task. Results show that PACE obtained higher recognition scores than ACE as found with CI r ecipients. Waldo Nogueira, Tamás Harczos, Bernd Edler, Jörn Ostermann, Andreas Büchner |
INTERSPEECH | 3 |
| 2006 | Wavelet Packet Filterbank for Speech Processing Strategies in Cochlear ImplantsabstractCurrent speech processing strategies for cochlear implants use a filterbank which decomposes the audio signals into multiple frequency bands each associated with one electrode. Pitch perception with cochlear implants is related to the number of electrodes inserted in the cochlea and to the rate of stimulation of these electrodes. The filterbank should, therefore, be able to analyze the time-frequency features of the audio signals while also exploiting the time-frequency features of the implant. This study investigates the influence on speech intelligibility in cochlear implant users when filterbanks with different time-frequency resolutions are used. Three filter-banks, based on the structure of a wavelet packet transform but using different basis functions, were designed. The filter-banks were incorporated into a commercial speech processing strategy and were tested on device users in an acute study. Waldo Nogueira, Andreas Giese, Bernd Edler, Andreas Büchner |
ICASSP (5) | 3 |
| 2005 | Motion-and aliasing-compensated prediction using a two-dimensional non-separable adaptive Wiener interpolation filterabstractIn the context of prediction with fractional-pel motion vector resolution it was shown, that aliasing components contained in an image signal are limiting the prediction accuracy obtained by motion compensation. In order to consider aliasing, quantisation and motion estimation errors, camera noise, etc., we analytically developed a two-dimensional (2D) non-separable interpolation filter, which is calculated for each frame independently by minimising the prediction error energy. For every fractional-pel position to be interpolated, an individual set of 2D filter coefficients is determined. As a result, a coding gain of up to 1,2 dB for HDTV-sequences and up to 0,5 dB for CIF-sequences compared to the standard H.264/AVC is obtained. Yuri Vatis, Bernd Edler, Dieu Thanh Nguyen, Jörn Ostermann |
ICIP (2) | 2 |
| 2002 | Sinusoidal coding using loudness-based component selectionabstractSinusoidal modelling forms the base of parametric audio coding systems, like MPEG-4 HILN, where it is combined with noise and transient models. A parametric encoder decomposes the audio signal into components that are described by appropriate models and represented by model parameters. To achieve efficient coding at very low bitrates, selection of the perceptually most relevant signal components (e.g. sinusoids) is essential, as only a limited number of component parameters can be conveyed in the bitstream. Various strategies for sinusoidal component selection have been proposed in the literature. This paper introduces a new, loudness-based strategy and tries to compare the different strategies using objective and subjective criteria. Heiko Purnhagen, Nikolaus Meine, Bernd Edler |
ICASSP | 3 |
| 2002 | Perceptual audio coding using adaptive pre- and post-filters and lossless compressionabstractThis paper proposes a versatile perceptual audio coding method that achieves high compression ratios and is capable of low encoding/decoding delay. It accommodates a variety of source signals (including both music and speech) with different sampling rates. It is based on separating irrelevance and redundancy reductions into independent functional units. This contrasts traditional audio coding where both are integrated within the same subband decomposition. The separation allows for the independent optimization of the irrelevance and redundancy reduction units. For both reductions, we rely on adaptive filtering and predictive coding as much as possible to minimize the delay. A psycho-acoustically controlled adaptive linear filter is used for the irrelevance reduction, and the redundancy reduction is carried out by a predictive lossless coding scheme, which is termed weighted cascaded least mean squared (WCLMS) method. Experiments are carried out on a database of moderate size which contains mono-signals of different sampling rates and varying nature (music, speech, or mixed). They show that the proposed WCLMS lossless coder outperforms other competing lossless coders in terms of compression ratios and delay, as applied to the pre-filtered signal. Moreover, a subjective listening test of the combined pre-filter/lossless coder and a state-of-the-art perceptual audio coder (PAC) shows that the new method achieves a comparable compression ratio and audio quality with a lower delay. Gerald Schuller, Bin Yu 0001, Dawei Huang, Bernd Edler |
IEEE Trans. Speech Audio Process. | 4 |
| 2000 | Audio coding using a psychoacoustic pre- and post-filterabstractA novel concept for perceptual audio coding is presented which is based on the combination of a pre- and post-filter, controlled by a psychoacoustic model, with a transform coding scheme. This paradigm allows modeling of the temporal and spectral shape of the masked threshold with a resolution independent of the used transform. By using frequency warping techniques the maximum possible detail for a given filter order can be made frequency-dependent and thus better adapted to the human auditory system. The filter coefficients are represented efficiently by LSF parameters which can be adaptively interpolated over time. First experiments with a system obtained by extending an existing transform codec showed that this approach can significantly improve the performance for speech signals, while the performance for other signals remained the same. Bernd Edler, Gerald Schuller |
ICASSP | 1 |
| 1997 | Tests on MPEG-4 audio codec proposals
Laura Contin, Bernd Edler, D. Meares, P. Schreiner |
Signal Process. Image Commun. | 2 |
| 1995 | Overlapping block transform: window design, fast algorithm, and an image coding experimentabstractA window design and fast algorithm for the overlapping block transform (OBT) of size N/spl times/L are presented. The presented algorithm for the OBT reduces the calculation complexity to an N/spl times/N transform with a fast algorithm and a simple preprocessing including windowing. A signal-independent window optimization strategy is introduced for image coding application. Results for a first-order Markov model and an image coding experiment show, that the coding gains of the optimized OBTs increase and blocking effects decrease with increasing window length L. A comparison with DCT-coding shows that the OBT, which has a slightly increased realization complexity, provides higher coding gain and a significant blocking effect reduction.> Miodrag R. Temerinac, Bernd Edler |
IEEE Trans. Commun. | 2 |
| 1993 | LINC: a common theory of transform and subband codingabstractA common theory of lapped orthogonal transforms (LOTs) and critically sampled filter banks, called L into N coding (LINC), is presented. The theory includes a unified analysis of both coding methods and identity relations between the transform, inverse transform, analysis filter bank, and synthesis filter bank. A design procedure for LINC analysis/synthesis systems, which satisfy the conditions for perfect reconstruction, is developed. The common LINC theory is used to define an ideal LINC system which is used, together with the power spectral density of the input signal, to calculate theoretical bounds for the coding gain. A generalized overlapping block transform (OBT) with time domain aliasing cancellation (TDAC) is used to approximate the ideal LINC. A generalization of the OBT includes multiple block overlap and additional windowing. A recursive design procedure for windows of arbitrary lengths is presented. The coding gain of the generalized OBT is higher than that of the Karhunen-Loeve transform (KLT) and close to the theoretical bounds for LINC. In the case of image coding, the generalized OBT reduces the blocking effects when compared with the DCT.> Miodrag R. Temerinac, Bernd Edler |
IEEE Trans. Commun. | 2 |
| 1992 | A unified approach to lapped orthogonal transformsabstractThe general conditions of exact reconstruction and a recursive design procedure for lapped orthogonal transform (LOT) with arbitrary length of overlapping are presented. It is shown that LOT can be realized with any standard block transform, discrete cosine transform (DCT), for example, and an additional processing. This processing must also satisfy the same conditions for exact reconfigurations and it may be pretransform processing in the time domain or post-transform processing in the transform domain. In a few examples it is shown that the LOT has a higher coding gain and smaller blocking effects then DCT. With the proposed LOT design procedure, two optimizations, the coding gain maximization and the blocking effect minimization, are presented and compared. Miodrag R. Temerinac, Bernd Edler |
IEEE Trans. Image Process. | 2 |