EDBT 2026 Demo / reviewers in the wild / expert
Bernd T. Meyer
dblp:31/9232
· DBLP profile ↗
59ranked-venue papers
11as first author
15since 2021 · last 2025
0000-0001-9190-2111ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 48 · 11 first-author · 10 since 2021Artificial intelligence and machine learning · 40 · 8 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Influence of Room Acoustics on Objective Voice Assessment Methods in the Context of Speech and Language Therapy
Sven Franz, Tanja Grewe, Bernd T. Meyer, Jörg Bitzer |
INTERSPEECH | 3 |
| 2025 | Hearing deficits of transformer-based ASR for anechoic and spatial signals
Dirk Eike Hoffner, Simon Weihe, Thomas Brand, Bernd T. Meyer |
INTERSPEECH | 4 |
| 2025 | Location-Aware Target Speaker Extraction for Hearing Aids
Daniel-José Alcala Padilla, Nils L. Westhausen, Swati Vivekananthan, Bernd T. Meyer |
INTERSPEECH | 4 |
| 2025 | Multilingual non-intrusive binaural intelligibility prediction based on phone classificationabstractSpeech intelligibility (SI) prediction models are a valuable tool for the development of speech processing algorithms for hearing aids or consumer electronics. For the use in realistic environments it is desirable that the SI model is non-intrusive (does not require separate input of original and degraded speech, transcripts or a-priori knowledge about the signals) and does a binaural processing of the audio signals. Most of the existing SI models do not fulfill all of these criteria. In this study, we propose an SI model based on phone probabilities obtained from a deep neural net. The model comprises a binaural enhancement stage for prediction of the speech recognition threshold (SRT) in realistic acoustic scenes. In the first part of the study, SRT predictions in different spatial configurations are compared to the results from normal-hearing listeners. On average, our approach produces lower errors and higher correlations compared to three intrusive baseline models. In the second part, we explore if measures relevant in spatial hearing, i.e., the intelligibility level difference (ILD) and the binaural ILD (BILD), can be predicted with our modeling approach. We also investigate if a language mismatch between training and testing the model plays a role when predicting ILD and BILD. This point is especially important for low-resource languages, where not thousands of hours of language material are available for training. Binaural benefits are predicted by our model with an error of 1.5 dB. This is slightly higher than the error with a competitive baseline MBSTOI (1.1 dB), but does not require separate input of original and degraded speech. We also find that good binaural predictions can be obtained with models that are not specifically trained with the target language. Jana Roßbach, Kirsten Wagener, Bernd T. Meyer |
Comput. Speech Lang. | 3 |
| 2025 | Non-intrusive binaural speech recognition prediction for hearing aid processingabstractHearing aids (HAs) often feature different signal processing algorithms to optimize speech recognition (SR) in a given acoustic environment. In this paper, we explore if models that predict SR performance of hearing-impaired (HI), aided users are applicable to automatically select the best algorithm. To this end, SR experiments are conducted with 19 HI subjects who are aided with an open-source HA. Listeners’ SR is measured in virtual, complex acoustic scenes with two distinct noise conditions using the different speech enhancement strategies implemented in this HA. For model-based selection, we apply a PHOneme-based Binaural Intelligibility model (PHOBI) based on our previous work and extended with a component for simulating hearing loss. The non-intrusive model utilizes a deep neural network to predict phone probabilities; the deterioration of these phone representations in the presence of noise or generally signal degradation is quantified and used as model output. PHOBI model is trained with 960 h of English speech signals, a broad range of noise signals and room impulse responses. The performance of model-based algorithm selection is measured with two metrics: (i) Its ability to rank the HA algorithms in the order of subjective SR results and (ii) the SR difference between the measured best algorithm and the model-based selection ( Δ SR). Results are compared to selections obtained with one non-intrusive and two intrusive models. PHOBI outperforms the non-intrusive and one of the intrusive models in both noise conditions, achieving significantly higher correlations ( r = 0 . 63 and 0.80). Δ SR scores are significantly lower (better) compared to the non-intrusive baseline (3.5% and 4.6% against 8.6% and 9.8%, respectively). The results in terms of Δ SR between PHOBI and the intrusive models are statistically not different, although PHOBI operates on the observed signal alone and does not require a clean reference signal. • A DNN-based model accurately predicts the hearing aid algorithm that optimizes speech recognition for its user. • Individual predictions are made for 19 hearing-impaired, aided users in complex acoustic scenes. • The DNN-based approach is non-intrusive and performs equally well as established, intrusive models for speech recognition prediction. Jana Roßbach, Nils L. Westhausen, Hendrik Kayser, Bernd T. Meyer |
Speech Commun. | 4 |
| 2024 | Joint prediction of subjective listening effort and speech intelligibility based on end-to-end learning
Dirk Eike Hoffner, Jana Roßbach, Bernd T. Meyer |
INTERSPEECH | 3 |
| 2024 | Real-Time Multichannel Deep Speech Enhancement in Hearing Aids: Comparing Monaural and Binaural Processing in Complex Acoustic ScenariosabstractDeep learning has the potential to enhance speech signals and increase their intelligibility for users of hearing aids. Deep models suited for real-world application should feature a low computational complexity and low processing delay of only a few milliseconds. In this paper, we explore deep speech enhancement that matches these requirements and contrast monaural and binaural processing algorithms in two complex acoustic scenes. Both algorithms are evaluated with objective metrics and in experiments with hearing-impaired listeners performing a speech-in-noise test. Results are compared to two traditional enhancement strategies, i.e., adaptive differential microphone processing and binaural beamforming. While in diffuse noise, all algorithms perform similarly, the binaural deep learning approach performs best in the presence of spatial interferers. Through a post-analysis, this can be attributed to improvements at low SNRs and to precise spatial filtering. Nils L. Westhausen, Hendrik Kayser, Theresa Jansen, Bernd T. Meyer |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Multilingual Query-by-Example Keyword Spotting with Metric Learning and Phoneme-to-Embedding MappingabstractIn this paper, we propose a multilingual query-by-example keyword spotting (KWS) system based on a residual neural network. The model is trained as a classifier on a multilingual keyword dataset extracted from Common Voice sentences and fine-tuned using circle loss. We demonstrate the generalization ability of the model to new languages and report a mean reduction in EER of 59.2% for previously seen and 47.9% for unseen languages compared to a competitive baseline. We show that the word embeddings learned by the KWS model can be accurately predicted from the phoneme sequences using a simple LSTM model. Our system achieves a promising accuracy for streaming keyword spotting and keyword search on Common Voice audio using just 5 examples per keyword. Experiments on the Hey-Snips dataset show a good performance with a false negative rate of 5.4% at only 0.1 false alarms per hour. Paul M. Reuter, Christian Rollwage, Bernd T. Meyer |
ICASSP | 3 |
| 2023 | Self-conducted speech audiometry using automatic speech recognition: Simulation results for listeners with hearing lossabstractSpeech-in-noise tests are an important tool for assessing hearing impairment, the successful fitting of hearing aids, as well as for research in psychoacoustics. An important drawback of many speech-based tests is the requirement of an expert to be present during the measurement, in order to assess the listener’s performance. This drawback may be largely overcome through the use of automatic speech recognition (ASR), which utilizes automatic response logging. However, such an unsupervised system may reduce the accuracy due to the introduction of potential errors. In this study, two different ASR systems are compared for automated testing: A system with a feed-forward deep neural network (DNN) from a previous study (Ooster et al., 2018), as well as a state-of-the-art system utilizing a time-delay neural network (TDNN). The dynamic measurement procedure of the speech intelligibility test was simulated considering the subjects’ hearing loss and selecting from real recordings of test participants. The ASR systems’ performance is investigated based on responses of 73 listeners, ranging from normal-hearing to severely hearing-impaired as well as read speech from cochlear implant listeners. The feed-forward DNN produced accurate testing results for NH and unaided HI listeners but a decreased measurement accuracy was found in the simulation of the adaptive measurement procedure when considering aided severely HI listeners, recorded in noisy environments with a loudspeaker setup. The TDNN system produces error rates of 0.6% and 3.0% for deletion and insertion errors, respectively. We estimate that the SRT deviation with this system is below 1.38 dB for 95% of the users. This result indicates that a robust unsupervised conduction of the matrix sentence test is possible with a similar accuracy as with a human supervisor even when considering noisy conditions and altered or disordered speech from elderly severely HI listeners and listeners with a CI. Jasper Ooster, Laura Tuschen, Bernd T. Meyer |
Comput. Speech Lang. | 3 |
| 2022 | Speech Intelligibility Prediction for Hearing-Impaired Listeners with the LEAP Modelabstract3498 Jana Roßbach, Rainer Huber, Saskia Röttges, Christopher F. Hauth, Thomas Biberger, Thomas Brand, Bernd T. Meyer, Jan Rennies |
INTERSPEECH | 7 |
| 2022 | tPLCnet: Real-time Deep Packet Loss Concealment in the Time Domain Using a Short Temporal Context
Nils L. Westhausen, Bernd T. Meyer |
INTERSPEECH | 2 |
| 2022 | Prediction of speech intelligibility with DNN-based performance measures
Angel Mario Castro Martinez, Constantin Spille, Jana Roßbach, Birger Kollmeier, Bernd T. Meyer |
Comput. Speech Lang. | 5 |
| 2021 | Non-Intrusive Binaural Prediction of Speech Intelligibility Based on Phoneme ClassificationabstractIn this study, we explore an approach for modeling speech intelligibility in spatial acoustic scenes. To this end, we combine a non-intrusive binaural frontend with a deep neural network (DNN) borrowed from a standard automatic speech recognition (ASR) system. The DNN estimates phoneme probabilities that degrade in the presence of noise and reverberation, which is quantified with an entropy-based measure. The model output is used to predict speech recognition thresholds, i.e., signal-to-noise ratio with 50% word recognition accuracy. It is compared to measured data obtained from eight normal-hearing listeners in acoustic scenarios with varying positions of localized maskers, different rooms and reverberation times. The model is non-intrusive; yet it produces a root mean squared error in the range of 0.6-2.1 dB, which is similar to results obtained with a reference model (0.3-1.8 dB) that uses oracle knowledge both in the frontend and in the backend stage. Jana Roßbach, Saskia Röttges, Christopher F. Hauth, Thomas Brand, Bernd T. Meyer |
ICASSP | 5 |
| 2021 | Acoustic Echo Cancellation with the Dual-Signal Transformation LSTM NetworkabstractThis paper applies the dual-signal transformation LSTM network (DTLN) to the task of real-time acoustic echo cancellation (AEC). The DTLN combines a short-time Fourier transform and a learned feature representation in a stacked network approach, which enables robust information processing in the time-frequency and in the time domain, which also includes phase information. The model is only trained on 60 h of real and synthetic echo scenarios. The training setup includes multi-lingual speech, data augmentation, additional noise and reverberation to create a model that should generalize well to a large variety of real-world conditions. The DTLN approach produces state-of-the-art performance on clean and noisy echo conditions reducing acoustic echo and additional noise robustly. The method outperforms the AEC-Challenge baseline by 0.30 in terms of Mean Opinion Score (MOS). Nils L. Westhausen, Bernd T. Meyer |
ICASSP | 2 |
| 2021 | Reduction of Subjective Listening Effort for TV Broadcast Signals With Recurrent Neural Networks
Nils L. Westhausen, Rainer Huber, Hannah Baumgartner, Ragini Sinha, Jan Rennies, Bernd T. Meyer |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2020 | DNN-Based Speech Presence Probability Estimation for Multi-Frame Single-Microphone Speech EnhancementabstractMulti-frame approaches for single-microphone speech enhancement, e.g., the multi-frame minimum-power-distortionless-response (MFMPDR) filter, are able to exploit speech correlations across neighboring time frames. In contrast to single-frame approaches such as the Wiener gain, it has been shown that multi-frame approaches achieve a substantial noise reduction with hardly any speech distortion, provided that an accurate estimate of the correlation matrices and especially the speech interframe correlation (IFC) vector is available. Typical estimation procedures of the IFC vector require an estimate of the speech presence probability (SPP) in each time-frequency (TF) bin. In this paper, we propose to use a bi-directional long short-term memory deep neural network (DNN) to estimate the SPP for each TF bin. Aiming at achieving a robust performance, the DNN is trained for various noise types and within a large signal-to-noise-ratio range. Experimental results show that the MFMPDR in combination with the proposed data-driven SPP estimator yields an increased speech quality compared to a state-of-the-art model-based SPP estimator. Furthermore, it is confirmed that exploiting interframe correlations in the MFMPDR is beneficial when compared to the Wiener gain especially in adverse scenarios. Marvin Tammen, Dörte Fischer, Bernd T. Meyer, Simon Doclo |
ICASSP | 3 |
| 2020 | Dual-Signal Transformation LSTM Network for Real-Time Noise SuppressionabstractThis paper introduces a dual-signal transformation LSTM network (DTLN) for real-time speech enhancement as part of the Deep Noise Suppression Challenge (DNS-Challenge). This approach combines a short-time Fourier transform (STFT) and a learned analysis and synthesis basis in a stacked-network approach with less than one million parameters. The model was trained on 500 h of noisy speech provided by the challenge organizers. The network is capable of real-time processing (one frame in, one frame out) and reaches competitive results. Combining these two types of signal transformations enables the DTLN to robustly extract information from magnitude spectra and incorporate phase information from the learned feature basis. The method shows state-of-the-art performance and outperforms the DNS-Challenge baseline by 0.24 points absolute in terms of the mean opinion score (MOS). Nils L. Westhausen, Bernd T. Meyer |
INTERSPEECH | 2 |
| 2019 | Improving Deep Models of Speech Quality Prediction through Voice Activity Detection and Entropy-based MeasuresabstractThis paper explores Deep machine listening for Estimating Speech Quality (DESQ), which predicts the perceived speech quality based on phoneme posterior probabilities obtained from a deep neural network. The degradation of phonemes is quantified with the entropy-based Gini measure that is compared to the mean temporal distance (MTD) proposed earlier. Since long speech pauses might have a large effect on the speech quality, we investigate if a voice activity detection (VAD) has a beneficial or detrimental effect on the predictive power of our model. The evaluation is performed by correlating the model output and mean opinion scores (MOS) of normal-hearing listeners who rated signals degraded by typical VoIP artifacts. While the Gini-based measure and MTD result in very similar predictions (with a lower computational cost for the Gini-measure), the VAD increases performance from r = 0.87 to r = 0.91 which is higher than three competing baselines (ITU-P.563, ANIQUE+, and SRM-Rnorm). Jasper Ooster, Bernd T. Meyer |
ICASSP | 2 |
| 2019 | "Computer, Test My Hearing": Accurate Speech Audiometry with Smart Speakers
Jasper Ooster, Pia Nancy Porysek Moreta, Jörg-Hendrik Bach, Inga Holube, Bernd T. Meyer |
INTERSPEECH | 5 |
| 2019 | DNN-based performance measures for predicting error rates in automatic speech recognition and optimizing hearing aid parameters
Angel Mario Castro Martinez, Lukas Gerlach 0001, Guillermo Payá-Vayá, Hynek Hermansky, Jasper Ooster, Bernd T. Meyer |
Speech Commun. | 6 |
| 2019 | Joint Estimation of Reverberation Time and Early-To-Late Reverberation Ratio From Single-Channel Speech SignalsabstractThe reverberation time (RT) and the early-to-late reverberation ratio (ELR) are two key parameters commonly used to characterize acoustic room environments. In contrast to conventional blind estimation methods that process the two parameters separately, we propose a model for joint estimation to predict the RT and the ELR simultaneously from single-channel speech signals from either full-band or sub-band frequency data, which is referred to as joint room parameter estimator (jROPE). An artificial neural network is employed to learn the mapping from acoustic observations to the RT and the ELR classes. Auditory-inspired acoustic features obtained by temporal modulation filtering of the speech time-frequency representations are used as input for the neural network. Based on an in-depth analysis of the dependency between the RT and the ELR, a two-dimensional (RT, ELR) distribution with constrained boundaries is derived, which is then exploited to evaluate four different configurations for jROPE. Experimental results show that-in comparison to the single-task ROPE system which individually estimates the RT or the ELR-jROPE provides improved results for both tasks in various reverberant and (diffuse) noisy environments. Among the four proposed joint types, the one incorporating multi-task learning with shared input and hidden layers yields the best estimation accuracies on average. When encountering extreme reverberant conditions with RTs and ELRs lying beyond the derived (RT, ELR) distribution, the type considering RT and ELR as a joint parameter performs robustly, in particular. From state-of-the-art algorithms that were tested in the acoustic characterization of environments challenge, jROPE achieves comparable results among the best for all individual tasks (RT and ELR estimation from full-band and sub-band signals). Feifei Xiong, Stefan Goetze, Birger Kollmeier, Bernd T. Meyer |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Prediction of Subjective Listening Effort from Acoustic Data with Non-Intrusive Deep ModelsabstractS.981-985 Paul Kranzusch, Rainer Huber, Melanie Krüger, Birger Kollmeier, Bernd T. Meyer |
INTERSPEECH | 5 |
| 2018 | Prediction of Perceived Speech Quality Using Deep Machine ListeningabstractS.976-980 Jasper Ooster, Rainer Huber, Bernd T. Meyer |
INTERSPEECH | 3 |
| 2018 | Predicting speech intelligibility with deep neural networks
Constantin Spille, Stephan Dieter Ewert, Birger Kollmeier, Bernd T. Meyer |
Comput. Speech Lang. | 4 |
| 2018 | Comparing human and automatic speech recognition in simple and complex acoustic scenes
Constantin Spille, Birger Kollmeier, Bernd T. Meyer |
Comput. Speech Lang. | 3 |
| 2018 | Evaluation of an automated speech-controlled listening test with spontaneous and read responses
Jasper Ooster, Rainer Huber, Birger Kollmeier, Bernd T. Meyer |
Speech Commun. | 4 |
| 2018 | Exploring Auditory-Inspired Acoustic Features for Room Acoustic Parameter Estimation From Monaural SpeechabstractRoom acoustic parameters that characterize acoustic environments can help to improve signal enhancement algorithms such as for dereverberation, or automatic speech recognition by adapting models to the current parameter set. The reverberation time (RT) and the early-to-late reverberation ratio (ELR) are two key parameters. In this paper, we propose a blind ROom Parameter Estimator (ROPE) based on an artificial neural network that learns the mapping to discrete ranges of the RT and the ELR from single-microphone speech signals. Auditory-inspired acoustic features are used as neural network input, which are generated by a temporal modulation filter bank applied to the speech time-frequency representation. ROPE performance is analyzed in various reverberant environments in both clean and noisy conditions for both fullband and subband RT and ELR estimations. The importance of specific temporal modulation frequencies is analyzed by evaluating the contribution of individual filters to the ROPE performance. Experimental results show that ROPE is robust against different variations caused by room impulse responses (measured versus simulated), mismatched noise levels, and speech variability reflected through different corpora. Compared to state-of-the-art algorithms that were tested in the acoustic characterisation of environments (ACE) challenge, the ROPE model is the only one that is among the best for all individual tasks (RT and ELR estimation from fullband and subband signals). Improved fullband estimations are even obtained by ROPE when integrating speech-related frequency subbands. Furthermore, the model requires the least computational resources with a real time factor that is at least two times faster than competing algorithms. Results are achieved with an average observation window of 3 s, which is important for real-time applications. Feifei Xiong, Stefan Goetze, Birger Kollmeier, Bernd T. Meyer |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Predicting error rates for unknown data in automatic speech recognitionabstractIn this paper we investigate methods to predict word error rates in automatic speech recognition in the presence of unknown noise types, which have not been seen during training. The performance measures operate on phoneme posteriorgrams that are obtained from neural nets. We compare average frame-wise entropy as a baseline measure to the mean temporal distance (M-Measure) and to the number of phonetic events. The latter is obtained by learning typical phoneme activations from clean training data, which are later applied as phoneme-specific matched filters to posteriorgrams (MaP). When exceeding a threshold after filtering, we register this as phonetic event. For test sets using 10 unknown noise types and a wide range of signal-to-noise ratios, we find M-Measure and MaP to produce predictions twice as accurate as the baseline measure. When excluding noise types that contain speech segments, a prediction error of 3.1% is achieved, compared to 15.0% for the baseline measure. Bernd T. Meyer, Sri Harish Reddy Mallidi, Hendrik Kayser, Hynek Hermansky |
ICASSP | 1 |
| 2017 | Combination strategy based on relative performance monitoring for multi-stream reverberant speech recognitionabstractA multi-stream framework with deep neural network (DNN) classifiers is applied to improve automatic speech recognition (ASR) in environments with different reverberation characteristics. We propose a room parameter estimation model to establish a reliable combination strategy which performs on either DNN posterior probabilities or word lattices. The model is implemented by training a multilayer perceptron incorporating auditory-inspired features in order to distinguish between and generalize to various reverberant conditions, and the model output is shown to be highly correlated to ASR performances between multiple streams, i.e., relative performance monitoring, in contrast to conventional mean temporal distance based performance monitoring for a single stream. Compared to traditional multi-condition training, average relative word error rate improvements of 7.7% and 9.4% have been achieved by the proposed combination strategies performing on posteriors and lattices, respectively, when the multi-stream ASR is tested in known and unknown simulated reverberant environments as well as realistically recorded conditions taken from REVERB Challenge evaluation set. Feifei Xiong, Stefan Goetze, Bernd T. Meyer |
ICASSP | 3 |
| 2017 | On DNN posterior probability combination in multi-stream speech recognition for reverberant environmentsabstractA multi-stream framework with deep neural network (DNN) classifiers has been applied in this paper to improve automatic speech recognition (ASR) performance in environments with different reverberation characteristics. We propose a room parameter estimation model to determine the stream weights for DNN posterior probability combination with the aim of obtaining reliable log-likelihoods for decoding. The model is implemented by training a multi-layer perceptron to distinguish between various reverberant environments. The method is tested in known and unknown environments against approaches based on inverse entropy and autoencoders, with average relative word error rate improvements of 46% and 29%, respectively, when performing multi-stream ASR in different reverberant situations. Feifei Xiong, Stefan Goetze, Bernd T. Meyer |
ICASSP | 3 |
| 2017 | Single-Ended Prediction of Listening Effort Based on Automatic Speech Recognition
Rainer Huber, Constantin Spille, Bernd T. Meyer |
INTERSPEECH | 3 |
| 2017 | Listening in the Dips: Comparing Relevant Features for Speech Recognition in Humans and Machines
Constantin Spille, Bernd T. Meyer |
INTERSPEECH | 2 |
| 2017 | On the relevance of auditory-based Gabor features for deep learning in robust speech recognition
Angel Mario Castro Martinez, Sri Harish Reddy Mallidi, Bernd T. Meyer |
Comput. Speech Lang. | 3 |
| 2017 | Combining Binaural and Cortical Features for Robust Speech RecognitionabstractThe segregation of concurrent speakers and other sound sources is an important ability of the human auditory system, but is missing in most current systems for automatic speech recognition (ASR), resulting in a large gap between human and machine performance. This study combines processing related to peripheral and cortical stages of the auditory pathway: A physiologically motivated binaural model estimates the positions of moving speakers to enhance the desired speech signal. Second, signals are converted to spectro-temporal Gabor features that resemble cortical speech representations and which have been shown to improve ASR in noisy conditions. Spectro-temporal Gabor features improve recognition results in all acoustic conditions under consideration compared with Mel-frequency cepstral coefficients. Binaural processing results in lower word error rates (WERs) in acoustic scenes with a concurrent speaker, whereas monaural processing should be preferred in the presence of a stationary masking noise. In-depth analysis of binaural processing identifies crucial processing steps such as localization of sound sources and estimation of the beamformer's noise coherence matrix, and shows how much each processing step affects the recognition performance in acoustic conditions with different complexity. Constantin Spille, Birger Kollmeier, Bernd T. Meyer |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Introducing Temporal Rate Coding for Speech in Cochlear Implants: A Microscopic Evaluation in Humans and Models
Anja Eichenauer, Mathias Dietz, Bernd T. Meyer, Tim Jürgens |
INTERSPEECH | 3 |
| 2016 | DNN-Based Automatic Speech Recognition as a Model for Human Phoneme Perception
Mats Exter, Bernd T. Meyer |
INTERSPEECH | 2 |
| 2016 | Neural Responses to Speech-Specific Modulations Derived from a Spectro-Temporal Filter Bank
Marina Frye, Cristiano Micheli, Inga M. Schepers, Gerwin Schalk, Jochem W. Rieger, Bernd T. Meyer |
INTERSPEECH | 6 |
| 2016 | Assessing Speech Quality in Speech-Aware Hearing Aids Based on Phoneme Posteriorgrams
Constantin Spille, Hendrik Kayser, Hynek Hermansky, Bernd T. Meyer |
INTERSPEECH | 4 |
| 2016 | Performance monitoring for automatic speech recognition in noisy multi-channel environmentsabstractIn many applications of machine listening it is useful to know how well an automatic speech recognition system will do before the actual recognition is performed. In this study we investigate different performance measures with the aim of predicting word error rates (WERs) in spatial acoustic scenes in which the type of noise, the signal-to-noise ratio, parameters for spatial filtering, and the amount of reverberation are varied. All measures under consideration are based on phoneme posteriorgrams obtained from a deep neural net. While frame-wise entropy exhibits only medium predictive power for factors other than additive noise, we found the medium temporal distance between posterior vectors (M-Measure) as well as matched phoneme filters (MaP) to exhibit excellent correlations with WER across all conditions. Since our results were obtained with simulated behind-the-ear hearing aid signals, we discuss possible applications for speech-aware hearing devices. Bernd T. Meyer, Sri Harish Reddy Mallidi, Angel Mario Castro Martinez, Guillermo Payá-Vayá, Hendrik Kayser, Hynek Hermansky |
SLT | 1 |
| 2015 | A study on joint beamforming and spectral enhancement for robust speech recognition in reverberant environmentsabstractThis work evaluates multi-microphone beamforming and single-microphone spectral enhancement strategies to alleviate the reverberation effect for robust automatic speech recognition (ASR) systems in different reverberant environments characterized by different reverberation times T60 and direct-to-reverberation ratios (DRRs). The systems consist of minimum variance distortionless response (MVDR) beamformers in combination with minimum mean square error (MMSE) estimators, and late reverberation spectral variance (LRSV) estimators, the latter employing a generalized model of the room impulse response (RIR). Various system architectures are analyzed with a focus on optimal speech recognition performance. The system combining an MVDR beamformer and a subsequent MMSE estimator was found to lead to the best results, with relative reductions of 27.7% compared to the baseline system. This is attributed to a more accurate LRSV estimate from spatial averaging and diffuse field refinement for the MMSE estimator. Feifei Xiong, Bernd T. Meyer, Stefan Goetze |
ICASSP | 2 |
| 2015 | Improving automatic speech recognition in spatially-aware hearing aidsabstractIn the context of ambient assisted living, automatic speech recognition (ASR) has the potential to provide textual support for hearing aid users in challenging acoustic conditions. In this paper we therefore investigate possibilities to improve ASR based on binaural hearing aid signals in complex acoustic scenes. Particularly, information about the spatial configuration of sound sources is exploited and estimated using a recently developed method that employs probabilistic information about the location of a target speaker (and a simultaneous localized masker) for robust real-time localization. Two different strategies are investigated: straightforward better-ear listening and a multi-channel beamforming system aiming at enhancement of a target speech source with additional suppression of localized masking sound. The latter method is also complemented by better-ear listening. Both approaches are evaluated in different acoustic scenarios containing moving target and interfering speakers or noise sources. Compared to using nonpreprocessed signals, we obtain average relative reductions in word error rate of 28.4% in the presence of a localized interfering noise, 19.2% in the case of a concurrent talker and 23.7% in presence of a concurrent talker in spatially diffuse noise. A post-analysis assesses the relation of localization performance and beamforming for improved speech recognition in complex acoustic scenes. Hendrik Kayser, Constantin Spille, Daniel Marquardt, Bernd T. Meyer |
INTERSPEECH | 4 |
| 2015 | Autonomous measurement of speech intelligibility utilizing automatic speech recognition
Bernd T. Meyer, Birger Kollmeier, Jasper Ooster |
INTERSPEECH | 1 |
| 2014 | Estimating room acoustic parameters for speech recognizer adaptation and combination in reverberant environmentsabstractThis work analyzes the influence of reverberation on automatic speech recognition (ASR) systems and how to compensate its influence, with special focus on the important acoustical parameters i.e. room reverberation time T60and clarity index C50. A multilayer perceptron (MLP) using features of a spectro-temporal filter bank as input is employed to identify the acoustic conditions spanning various reverberant scenarios. The posterior probabilities of the MLP are used to design a novel selection scheme for adaptation in a cluster-based manner and for system combination achieved by recognizer output voting error reduction (ROVER). A comparison of word error rates is performed considering different training modes, and an average relative improvement of 7.1% is obtained by the proposed system compared to conventional multistyle training. Feifei Xiong, Stefan Goetze, Bernd T. Meyer |
ICASSP | 3 |
| 2014 | Should deep neural nets have ears? the role of auditory features in deep learning approachesabstractFeatures inspired by the auditory system have previously demonstrated improvement in automatic speech recognition (ASR). Similarly, the use of Deep Neural Networks (DNN) was found to outperform classic approaches to ASR in many conditions. Since DNNs have the potential to learn the task relevant features from a conventional filter bank output, we investigate if the combination of auditory features and deep learning should be preferred over self-learned patterns. Specifically, noise-robust Gabor features and Amplitude Modulation Filter-Bank (AMFB) features, highly invariant against reverberation, are used as input to a state-of-the-art ASR system incorporating DNN processing. On the Aurora-4 task, both mel-frequency cepstral coefficients (MFCC) and filter bank (FBank) features are outperformed in many acoustic conditions through auditory processing, yielding average relative improvements of up to 69% over MFCC and 21% over the commonly used DNN-FBank setup. This highlights the mutual benefit of auditory signal processing and recent advances in machine learning. Angel Mario Castro Martinez, Niko Moritz, Bernd T. Meyer |
INTERSPEECH | 3 |
| 2014 | Identifying the human-machine differences in complex binaural scenes: what can be learned from our auditory systemabstractPrevious comparisons of human speech recognition (HSR) and automatic speech recognition (ASR) focused on monaural signals in additive noise, and showed that HSR is far more robust against intrinsic and extrinsic sources of variation than conventional ASR. The aim of this study is to analyze the man-machine gap (and its causes) in more complex acoustic scenarios, particularly in scenes with two moving speakers, reverberation and diffuse noise. Responses of nine normal-hearing listeners are compared to errors of an ASR system that employs a binaural model for direction-of-arrival estimation and beamforming for signal enhancement. The overall man-machine gap is measured in terms for the speech recognition threshold (SRT), i.e., the signal-to-noise ratio at which a 50 % recognition rate is obtained. The comparison shows that the gap amounts to 16.7 dB SRT difference which exceeds the difference of 10 dB found in monaural situations. Based on cross comparisons that use oracle knowledge (e.g., the speakers’ true position), incorrect responses are attributed to localization errors (7 dB) or missing spectral information to distinguish between speakers with different gender (3 dB). The comparison hence identifies specific ASR components that can profit from learning from binaural auditory signal processing. Constantin Spille, Bernd T. Meyer |
INTERSPEECH | 2 |
| 2013 | Spectro-temporal features for noise-robust speech recognition using power-law nonlinearity and power-bias subtractionabstractPrevious work has demonstrated that spectro-temporal Gabor features reduced word error rates for automatic speech recognition under noisy conditions. However, the features based on mel spectra were easily corrupted in the presence of noise or channel distortion. We have exploited an algorithm for power normalized cepstral coefficients (PNCCs) to generate a more robust spectro-temporal representation. We refer to it as power normalized spectrum (PNS), and to the corresponding output processed by Gabor filters and MLP nonlinear weighting as PNS-Gabor. We show that the proposed feature outperforms state-of-the-art noise-robust features, ETSI-AFE and PNCC for both Aurora2 and a noisy version of the Wall Street Jounal (WSJ) corpus. A comparison of the individual processing steps of mel spectra and PNS shows that power bias subtraction is the most important aspect of PNS-Gabor features to provide an improvement over Mel-Gabor features. The result indicates that Gabor processing compensates the limitation of PNCC for channels with frequency-shift characteristic. Overall, PNS-Gabor features decrease the word error rate by 32% relative to MFCC and 13% relative to PNCC in Aurora2. For noisy WSJ, they decrease the word error rate by 30.9% relative to MFCC and 24.7% relative to PNCC. Shuo-Yiin Chang, Bernd T. Meyer, Nelson Morgan |
ICASSP | 2 |
| 2013 | Using binarual processing for automatic speech recognition in multi-talker scenesabstractThe segregation of concurrent speakers and other sound sources is an important aspect of the human auditory system but is missing in most current systems for automatic speech recognition (ASR), resulting in a large gap between human and machine performance. The present study uses a physiologically-motivated model of binaural hearing to estimate the position of moving speakers in a noisy environment by combining methods from Computational Auditory Scene Analysis (CASA) and ASR. The binaural model is paired with a particle filter and a beamformer to enhance spoken sentences that are transcribed by the ASR system. Results based on an evaluation in clean, anechoic two-speaker condition shows the word recognition rates to be increased from 30.8% to 72.6%, demonstrating the potential of the CASA-based approach. In different noisy environments, improvements were also observed for SNRs of 5 dB and above, which was attributed to the average tracking errors that were consistent over a wide range of SNRs. Constantin Spille, Mathias Dietz, Volker Hohmann, Bernd T. Meyer |
ICASSP | 4 |
| 2013 | Blind estimation of reverberation time based on spectro-temporal modulation filteringabstractA novel method for blind estimation of the reverberation time (RT60) is proposed based on applying spectro-temporal modulation filters to time-frequency representations. 2D-Gabor filters arranged in a filterbank enable an analysis of the properties of temporal, spectral, and spectro-temporal filtering for this task. Features are used as input to a multi-layer perceptron (MLP) classifier combined with a simple decision rule that attributes a specific RT60 to a given utterance and allows to assess the reliability of the approach for different resolutions of RT60 classification. While the filter set including temporal, spectral, and spectro-temporal filters already outperforms an MFCC baseline, the error rates are further reduced when relying on diagonal spectro-temporal filters alone. The average error rate is 1.9% for the best feature set, which corresponds to a relative reduction of 58.3% compared to the MFCC baseline for RT60s in 0.1 s resolution. Feifei Xiong, Stefan Goetze, Bernd T. Meyer |
ICASSP | 3 |
| 2013 | What's the difference? comparing humans and machines on the Aurora 2 speech recognition task
Bernd T. Meyer |
INTERSPEECH | 1 |
| 2012 | Spectro-temporal Gabor features for speaker recognitionabstractIn this work, we have investigated the performance of 2D Gabor features (known as spectro-temporal features) for speaker recognition. Gabor features have been used mainly for automatic speech recognition (ASR), where they have yielded improvements. We explored different Gabor feature implementations, along with different speaker recognition approaches, on ROSSI [1] and NIST SRE08 databases. Using the noisy ROSSI database, the Gabor features performed as well as the MFCC features standalone, and score-level combination of Gabor and MFCC features resulted in an 8% relative EER improvement over MFCC features standalone. These results demonstrated the value of both spectral and temporal information for feature extraction, and the complementarity of Gabor features to MFCC features. Howard Lei, Bernd T. Meyer, Nikki Mirghafori |
ICASSP | 2 |
| 2012 | Hooking up spectro-temporal filters with auditory-inspired representations for robust automatic speech recognitionabstractSpectro-temporal filtering has been shown to result in features that can help to increase the robustness of automatic speech recognition (ASR) in the past. We replace the spectro-temporal representation used in previous work with spectrograms that incorporate knowledge about the signal processing of the human auditory system and which are derived from Power-Normalized Cepstral Coefficients (PNCCs). 2D-Gabor filters are applied to these spectrograms to extract features evaluated on a noisy digit recognition task. The filter bank is adapted to the new representation by optimizing the spectral modulation frequencies associated with each Gabor function. A comparison of optimized parameters and the spectral modulation of vowels shows a good match between optimized and expected range of frequencies. When processed with a non-linear neural net and combined with PNCCs, Gabor features decrease the error rate compared to the baseline and PNCCs by at least 19%. Bernd T. Meyer, Constantin Spille, Birger Kollmeier, Nelson Morgan |
INTERSPEECH | 1 |
| 2011 | Comparing Different Flavors of Spectro-Temporal Features for ASRabstractIn the last decade, several studies have shown that the robustness of ASR systems can be increased when 2D Gabor filters are used to extract specific modulation frequencies from the input pattern. This paper analyzes important design parameters for spectro-temporal features based on a Gabor filter bank: We perform experiments with filters that exhibit different phase sensitivity. Further, we analyze if non-linear weighting with a multi-layer perceptron (MLP) and a subsequent concatenation with mel-frequency cepstral coefficients (MFCCs) has beneficial effects. For the Aurora2 noisy digit recognition task, the use of phase sensitive filters improved the MFCC baseline, whereas using filters that neglect phase information did not. While MLP processing alone did not have a large effect on the overall performance, the best results were obtained for MLP-processed phase sensitive filters and added MFCCs, with relative error reductions of over 40% for both noisy and clean training. Bernd T. Meyer, Suman V. Ravuri, Marc René Schädler, Nelson Morgan |
INTERSPEECH | 1 |
| 2011 | Robustness of spectro-temporal features against intrinsic and extrinsic variations in automatic speech recognition
Bernd T. Meyer, Birger Kollmeier |
Speech Commun. | 1 |
| 2010 | Learning from human errors: prediction of phoneme confusions based on modified ASR trainingabstractIn an attempt to improve models of human perception, the recognition of phonemes in nonsense utterances was predicted with automatic speech recognition (ASR) in order to analyze its applicability for modeling human speech recognition (HSR) in noise. In the first experiments, several feature types are used as input for an ASR system; the resulting phoneme scores are compared to listening experiments using the same speech data. With conventional training, the highest correlation between predicted and measured recognition was observed for perceptual linear prediction features (r = 0.84). Secondly, a new training paradigm for ASR is proposed with the aim of improving the prediction of phoneme intelligibility. For this perceptual training, the original utterance labels are modified based on the confusions measured in HSR tests. The modified ASR training improved the overall prediction, with the best models (r = 0.89) exceeding those obtained with conventional training Bernd T. Meyer, Birger Kollmeier |
INTERSPEECH | 1 |
| 2009 | Complementarity of MFCC, PLP and Gabor features in the presence of speech-intrinsic variabilitiesabstractIn this study, the effect of speech-intrinsic variabilities such as speaking rate, effort and speaking style on automatic speech recognition (ASR) is investigated. We analyze the influence of such variabilities as well as extrinsic factors (i.e., additive noise) on the most common features in ASR (mel-frequency cepstral coefficients and perceptual linear prediction features) and spectro-temporal Gabor features. MFCCs performed best for clean speech, whereas Gabors were found to be the most robust feature in extrinsic variabilities. Intrinsic variations were found to have a strong impact on error rates. While performance with MFCCs and PLPs was degraded in much the same way, Gabor features exhibit a different sensivity towards these variabilities and are, e.g., well-suited to recognize speech with varying pitch. The results suggest that spectro-temporal and classic features carry complementary information, which could be exploited in feature-stream experiments. Index Terms: automatic speech recognition, speech-intrinsic variabilities, feature extraction, spectro-temporal features Bernd T. Meyer, Birger Kollmeier |
INTERSPEECH | 1 |
| 2008 | The non-native consonant challenge for european languagesabstractThis paper reports on a multilingual investigation into the effects of different masker types on native and non-native perception in a VCV consonant recognition task. Native listeners outperformed 7 other language groups, but all groups showed a similar ranking of maskers. Strong first language (L1) interference was observed, both from the sound system and from the L1 orthography. Universal acoustic-perceptual tendencies are also at work in both native and non-native sound identifications in noise. The effect of linguistic distance, however, was less clear: in large multilingual studies, listener variables may overpower other factors. María Luisa García Lecumberri, Martin Cooke, Francesco Cutugno, Mircea Giurgiu, Bernd T. Meyer, Odette Scharenborg, Wim A. van Dommelen, Jan Volín |
INTERSPEECH | 5 |
| 2008 | Optimization and evaluation of Gabor feature sets for ASRabstractIn order to enhance automatic speech recognition performance in adverse conditions, Gabor features motivated by physiolog-ical measurements in the primary auditory cortex were opti-mized and evaluated. In the Aurora 2 experimental setup such localized, spectro-temporal filters combined with a Tandem sys-tem yield robust performance with a feature set size of 30. Im-proved results can be obtained when using a Hanning window instead of a cut-off Gaussian envelope due to better modulation frequency characteristics. An analysis of complementarity of Gabor and MFCC features shows that errors could be reduced by 55 % with a perfect classifier. In a real world scenario, a rela-tive WER reduction of 15 % compared to a competitive baseline is achieved by combining the feature types, indicating the po-tential of this class of physiologically motivated features. Index Terms: spectro temporal features, automatic speech recognition, Gabor features 1. Bernd T. Meyer, Birger Kollmeier |
INTERSPEECH | 1 |
| 2007 | Phoneme confusions in human and automatic speech recognitionabstractA comparison between automatic speech recognition (ASR) and human speech recognition (HSR) is performed as prerequisite for identifying sources of errors and improving feature extraction in ASR. HSR and ASR experiments are carried out with the same logatome database which consists of nonsense syllables. Two different kinds of signals are presented to human listeners: First, noisy speech samples are converted to Mel-frequency cepstral coefficients which are resynthesized to speech, with information about voicing and fundamental frequency being discarded. Second, the original signals with added noise are presented, which is used to evaluate the loss of information caused by the process of resynthesis. The analysis also covers the degradation of ASR caused by dialect or accent and shows that different error patterns emerge for ASR and HSR. The information loss induced by the calculation of ASR features has the same effect as a deteriation of the SNR by 10 dB. Index Terms: human speech recognition, automatic speech recognition, dialect, accent, phoneme confusions, MFCC Bernd T. Meyer, Matthias Wächter, Thomas Brand, Birger Kollmeier |
INTERSPEECH | 1 |
| 2005 | Oldenburg logatome speech corpus (OLLO) for speech recognition experiments with humans and machinesabstractThis paper introduces the new OLdenburg LOgatome speech corpus (OLLO) and outlines design considerations during its creation. OLLO is distinct from previous ASR corpora as it specifically targets (1) the fair comparison between human and machine speech recognition performance, and (2) the realistic representation of intrinsic variabilities in speech that are significant for automatic speech recognition (ASR) systems. To enable an unbiased human-machine comparison, OLLO is designed for recognition of individual phonemes that are embedded in logatomes, specifically, three-phoneme sequences with no semantic information. A balanced set of target-phonemes important for human and automatic speech recognition has been chosen, drawing on pilot ASR studies and cross-fertilization from the field of human speech intelligibility testing. Several intrinsic variabilities in speech are represented in OLLO, by recording from 40 speakers from four German dialect regions, and by covering six articulation characteristics. Results from preliminary phonetic time-labeling and ASR experiments are promising and consistent with corpus variabilities. 1. Thorsten Wesker, Bernd T. Meyer, Kirsten Wagener, Jörn Anemüller, Alfred Mertins, Birger Kollmeier |
INTERSPEECH | 2 |