EDBT 2026 Demo / reviewers in the wild / expert
Stefan Goetze
dblp:63/5923
· DBLP profile ↗
42ranked-venue papers
1as first author
18since 2021 · last 2026
0000-0003-1044-7343ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 1 first-author · 17 since 2021Artificial intelligence and machine learning · 16 · 7 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards automating the Frenchay dysarthria assessment: Can neural phoneme posteriorgrams inform the analysis of dysarthric speech?abstractDysarthria is a type of motor speech disorder that reflects abnormalities in motor movements required for speech production. In clinical practice, identifying characteristic signs and symptoms of the neuropathophysiology underlying a dysarthria is vital for diagnosis and management. The gold standard for dysarthria assessment is auditory-perceptual evaluation by a speech and language therapist for differential diagnosis and management decisions. As the process is time-consuming for clinicians, there is growing interest in automatic dysarthria assessment (ADA). Recent approaches to ADA primarily focus on the classification of broad intelligibility or speech severity labels. However, this does not have much clinical utility and the assessment of communication-relevant parameters do not distinguish between dysarthria types and pathomechanisms. Studies on the classification of dysarthria function or clinical test protocol scores focusing on aspects of dysarthric speech production (such as the Frenchay dysarthria assessment (FDA)) are limited. Therefore, this paper focuses on the preliminary steps towards clinically interpretable ADA, including automatic FDA assessment. The phoneme posteriorgram (PPG) is a time-varying categorical distribution over acoustic speech units, and recent work demonstrates interpretable speech pronunciation distance for downstream tasks, e.g. pronunciation reconstruction. This work extends recent advances in posterior-based phoneme research and mispronunciation models to dysarthria assessment, exploring the extent to which dysarthric speech features in the FDA (identified by auditory-perceptual evaluation in clinical practice) are captured by PPG information. To achieve this, FDA aspects are systematically evaluated. The results show that interpretable PPG probability can capture dysarthric speech features that are related to motor system dysfunction. Wing-Zin Leung, Heidi Christensen, Stefan Goetze |
Speech Commun. | 3 |
| 2025 | Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric RecordingsabstractAudiovisual active speaker detection (ASD) addresses the task of determining the speech activity of a candidate speaker given acoustic and visual data. Typically, systems model the temporal correspondence of audiovisual cues, such as the synchronisation between speech and lip movement. Recent work has explored extending this paradigm by additionally leveraging speaker embeddings extracted from candidate speaker reference speech. This paper proposes the speaker comparison auxiliary network (SCAN) which uses speaker-specific information from both reference speech and the candidate audio signal to disambiguate challenging scenes when the visual signal is unresolvable. Furthermore, an improved method for enrolling face-speaker libraries is developed, which implements a self-supervised approach to video-based face recognition. Fitting with the recent proliferation of wearable devices, this work focuses on improving speaker-embedding-informed ASD in the context of egocentric recordings, which can be characterised by acoustic noise and highly dynamic scenes. SCAN is implemented with two well-established baselines, namely TalkNet and Light-ASD; yielding a relative improvement in mAP of 14.5% and 10.3% on the Ego4D benchmark, respectively. Jason Clarke, Yoshihiko Gotoh, Stefan Goetze |
ICASSP | 3 |
| 2025 | MetricGAN+KAN: Kolmogorov-Arnold Networks in Metric-Driven Speech Enhancement SystemsabstractNeural-network-based speech enhancement (SE) approaches have shown to be particularly powerful in combination with perceptually motivated metrics to produce high-quality enhanced speech signals. Among these deep learning (DL)-based SE models, MetricGAN and its extension can generate output signals directly optimising quality metrics. The recently proposed Kolmogorov-Arnold networks (KANs) with learnable activation functions have shown great success in replacing multi-layer perceptrons (MLPs). This work proposes the use of KANs in a MetricGAN framework and analyses their performance in replacing different types of network layers. The best-performing proposed MetricGAN+KAN model uses approximately 80% fewer parameters and achieves 13.2% higher SE performance (measured by PESQ) on the Voicebank-DEMAND dataset, compared to the MetricGAN+ baseline. Yemin Mai, Stefan Goetze |
ICASSP | 2 |
| 2025 | Alzheimer's Dementia Detection Using Perplexity from Paired Large Language Models
Heidi Christensen, Stefan Goetze |
INTERSPEECH | 3 |
| 2024 | Multi-CMGAN+/+: Leveraging Multi-Objective Speech Quality Metric Prediction for Speech EnhancementabstractNeural network based approaches to speech enhancement have shown to be particularly powerful, being able to leverage a data-driven approach to result in a significant performance gain versus other approaches. Such approaches are reliant on artificially created labelled training data such that the neural model can be trained using intrusive loss functions which compare the output of the model with clean reference speech. Performance of such systems when enhancing real-world audio often suffers relative to their performance on simulated test data. In this work, a non-intrusive multi-metric prediction approach is introduced, wherein a model trained on artificial labelled data using inference of an adversarially trained metric prediction neural network. The proposed approach shows improved performance versus state-of-the-art systems on the recent CHiME-7 challenge unsupervised domain adaptation speech enhancement (UDASE) task evaluation sets. Index Terms: speech enhancement, model generalisation, generative adversarial networks, conformer, metric prediction George Close, William Ravenscroft, Thomas Hain, Stefan Goetze |
ICASSP | 4 |
| 2024 | Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users Using Intermediate ASR Features and Human Memory ModelsabstractNeural networks have been successfully used for non-intrusive speech intelligibility prediction. Recently, the use of feature representations sourced from intermediate layers of pre-trained self-supervised and weakly-supervised models has been found to be particularly useful for this task. This work combines the use of Whisper ASR decoder layer representations as neural network input features with an exemplar-based, psychologically motivated model of human memory to predict human intelligibility ratings for hearing-aid users. Substantial performance improvement over an established intrusive HASPI baseline system is found, including on enhancement systems and listeners unseen in the training data, with a root mean squared error of 25.3 compared with the baseline of 28.7. Rhiannon Mogridge, George Close, Robert Sutherland, Thomas Hain, Jon Barker, Stefan Goetze, Anton Ragni |
ICASSP | 6 |
| 2024 | Combining Conformer and Dual-Path-Transformer Networks for Single Channel Noisy Reverberant Speech SeparationabstractSeparation of overlapping speakers remains an active area of speech technology research. Many deep neural network (DNN) separation models propose modelling local and global temporal context separately using alternating DNN layers. Two such models are SepFormer and TD-Conformer. The largest configurations of each have comparable computational cost and similar performance; with SepFormer performing better on anechoic data and TD-Conformer yielding better results on noisy reverberant data. This work combines these two model types to gain insights into how their computational characteristics affect their performance. The generalization benefits of the larger model size of the conformer layers are demonstrated both on the WHAMR and the out-of-domain far-field evaluation set MC-WSJ-AV across a number of evaluation metrics. The proposed model is able to achieve 22.1 dB and 14.7 dB average scale-invariant signal-to-distortion ratio (SISDR) improvement when trained and evaluated on WSJ0-2Mix and WHAMR, respectively. The model trained using WHAMR is able to achieve 4.3 dB average SISDR improvement on the out-of-domain MC-WSJ-AV dataset. William Ravenscroft, Stefan Goetze, Thomas Hain |
ICASSP | 2 |
| 2024 | Active Learning for Sound Event Classification Using Bayesian Neural Networks with Gaussian Variational PosteriorabstractManual annotation of audio material is cumbersome. Active learning aims at minimizing the annotation effort by iteratively selecting an acquisition batch of unlabeled data, asking a human to annotate the selected data and re-training a classifier until an annotation budget is depleted. In this paper we propose the Gaussian-dense active learning (GDAL) algorithm to train a sound event classifier. The classifier is a Bayesian neural network where the weights are normally distributed. This is in contrast to conventional neural networks where weights are not distributed, but have assigned values. The Bayesian nature of the classifier empowers GDAL to select acquisition batches from a set of unlabeled audio clips based on their estimated informativeness. Evaluation results on the UrbanSound8k dataset show that GDAL outperforms a state-of-the-art algorithm based on medoid active learning for all considered annotation budgets and an algorithm based on dropout active learning for sufficiently large annotation budgets. Stepan Shishkin, Danilo Hollosi, Stefan Goetze, Simon Doclo |
ICASSP | 3 |
| 2024 | Refining Text Input For Augmentative and Alternative Communication (AAC) Devices: Analysing Language Model Layers For OptimisationabstractCommunication impairments are prevalent among a significant proportion of individuals. Methods of Augmentative and Alternative Communication (AAC) can support people with speech disorders (PwSD) to some extent, but AAC users encounter substantial difficulties when engaging in open-domain social interactions, especially involving multiple participants. This is mainly due to the significant communication rate gap between typical speakers and AAC users. Large Language Models (LLM) offer a solution by providing predictions of the next words or sentences. This work analyses refining the prediction capabilities of Masked Language Models (MLM) for AAC users by performing layer-wise analysis specifically for word prediction on an AAC corpus. Experiments show that fine-tuning only specific low-performing LLM layers leads to better results than fine-tuning of the entire model. Fine-tuning of specific layers of a Robust Bidirectional Encoder Representations from Transformers (RoBERTa) model outperforms other tested models; for qualitative evaluation and informal prototype AAC device testing. Fine-tuning the word predictions in an AAC context results in approx. 20% increase in average communication rate (across different communication scenarios) to input speed of approx. 30 words per minute (WPM). Hussein Yusufali, Roger K. Moore, Stefan Goetze |
ICASSP | 3 |
| 2024 | CADGE: Context-Aware Dialogue Generation Enhanced with Graph-Structured Knowledge AggregationabstractCommonsense knowledge is crucial to many natural language processing tasks.Existing works usually incorporate graph knowledge with conventional graph neural networks (GNNs), resulting in a sequential pipeline that compartmentalizes the encoding processes for textual and graph-based knowledge.This compartmentalization does, however, not fully exploit the contextual interplay between these two types of input knowledge.In this paper, a novel context-aware graph-attention model (Context-aware GAT) is proposed, designed to effectively assimilate global features from relevant knowledge graphs through a contextenhanced knowledge aggregation mechanism.Specifically, the proposed framework employs an innovative approach to representation learning that harmonizes heterogeneous features by amalgamating flattened graph knowledge with text data.The hierarchical application of graph knowledge aggregation within connected subgraphs, complemented by contextual information, to bolster the generation of commonsensedriven dialogues is analyzed.Empirical results demonstrate that our framework outperforms conventional GNN-based language models in terms of performance.Both, automated and human evaluations affirm the significant performance enhancements achieved by our proposed model over the concept flow baseline. Related WorkRecently, much work has focused on augmenting dialogue systems with additional background knowledge.Such works can be divided into dia- Tyler Loakman, Bohao Yang, Stefan Goetze, Chenghua Lin 0002 |
INLG | 5 |
| 2024 | Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech SynthesisabstractThis is a repository copy of Training data augmentation for dysarthric automatic speech recognition by text-to-dysarthric-speech synthesis. Wing-Zin Leung, Mattias Cross, Anton Ragni, Stefan Goetze |
INTERSPEECH | 4 |
| 2024 | Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech RecognitionabstractOne solution to automatic speech recognition (ASR) of overlapping speakers is to separate speech and then perform ASR on the separated signals.Commonly, the separator produces artefacts which often degrade ASR performance.Addressing this issue typically requires reference transcriptions to jointly train the separation and ASR networks.This is often not viable for training on real-world in-domain audio where reference transcript information is not always available.This paper proposes a transcription-free method for joint training using only audio signals.The proposed method uses embedding differences of pre-trained ASR encoders as a loss with a proposed modification to permutation invariant training (PIT) called guided PIT (GPIT).The method achieves a 6.4% improvement in word error rate (WER) measures over a signal-level loss and also shows enhancement improvements in perceptual measures such as short-time objective intelligibility (STOI). William Ravenscroft, George Close, Stefan Goetze, Thomas Hain, Mohammad Soleymanpour, Anurag Chowdhury, Mark C. Fuhs |
INTERSPEECH | 3 |
| 2023 | Improving Audiovisual Active Speaker Detection in Egocentric Recordings with the Data-Efficient Image TransformerabstractFuture augmented reality devices have the capacity to enhance human perception and provide assistive functions in complex communication scenarios. Active speaker detection (ASD) systems that are robust to egocentric data are critical to this. Egocentric ASD is challenging due to overlapping speech, single-channel recording, and dynamic scenes. A novel module that uses a data-efficient image transformer (DeiT) to extract features encapsulating the acoustic properties of each scene, and a positional conditioning mechanism is proposed. The module is evaluated in conjunction with TalkNet, an existing ASD architecture, on two audiovisual datasets: Ego4D (egocentric) and AVA-ActiveSpeaker (exocentric), achieving 29% and 0.38% relative improvement in mean Average Precision (mAP), respectively, while retaining a parameter efficient build. A qualitative analysis is also presented, implicitly demonstrating that contextual information is leveraged. Jason Clarke, Yoshihiko Gotoh, Stefan Goetze |
ASRU | 3 |
| 2023 | On Time Domain Conformer Models for Monaural Speech Separation in Noisy Reverberant Acoustic EnvironmentsabstractSpeech separation remains an important topic for multispeaker technology researchers. Convolution augmented transformers (conformers) have performed well for many speech processing tasks but have been under-researched for speech separation. Most recent state-of-the-art (SOTA) separation models have been time-domain audio separation networks (TasNets). A number of successful models have made use of dual-path (DP) networks which sequentially process local and global information. Time domain conformers (TD-Conformers) are an analogue of the DP approach in that they also process local and global context sequentially but have a different time complexity function. It is shown that for realistic shorter signal lengths, conformers are more efficient when controlling for feature dimension. Subsampling layers are proposed to further improve computational efficiency. The best TD-Conformer achieves $14.6 \mathrm{~dB}$ and $21.2 \mathrm{~dB}$ SISDR improvement on the WHAMR and WSJO2Mix benchmarks, respectively. William Ravenscroft, Stefan Goetze, Thomas Hain |
ASRU | 2 |
| 2023 | Perceive and Predict: Self-Supervised Speech Representation Based Loss Functions for Speech EnhancementabstractRecent work in the domain of speech enhancement has explored the use of self-supervised speech representations to aid in the training of neural speech enhancement models. However, much of this work focuses on using the deepest or final outputs of self supervised speech representation models, rather than the earlier feature encodings. The use of self supervised representations in such a way is often not fully motivated. In this work it is shown that the distance between the feature encodings of clean and noisy speech correlate strongly with psychoacoustically motivated measures of speech quality and intelligibility, as well as with human Mean Opinion Score (MOS) ratings. Experiments using this distance as a loss function are performed and improved performance over the use of STFT spectrogram distance based loss as well as other common loss functions from speech enhancement literature is demonstrated using objective measures such as perceptual evaluation of speech quality (PESQ) and short-time objective intelligibility (STOI). George Close, William Ravenscroft, Thomas Hain, Stefan Goetze |
ICASSP | 4 |
| 2023 | Moving Towards Non-Binary Gender Identification Via Analysis of System Errors in Binary Gender ClassificationabstractThis paper aims to analyse human perceptions of gender in speech signals, focusing on signals that are misclassified by methods for binary gender classification, looking at the features of speech signals that are more likely to be misclassified, or classified as either nonbinary or unclassifiable. The paper also analyses how human subjects perform in classifying such speech signals to gain insight into differences between machine and human performance levels. It is shown that gender classification systems and human ratings lack inter-annotator agreement, as do human ratings considered individually. There is also discussion of the suitability of continuing to use a binary system for gender in the field. This work fits into a larger body of research ongoing in the area of speech technology for trans-gender voice therapy. Sebastian Ellis, Stefan Goetze, Heidi Christensen |
ICASSP | 2 |
| 2023 | Deformable Temporal Convolutional Networks for Monaural Noisy Reverberant Speech SeparationabstractSpeech separation models are used for isolating individual speakers in many speech processing applications. Deep learning models have been shown to lead to state-of-the-art (SOTA) results on a number of speech separation benchmarks. One such class of models known as temporal convolutional networks (TCNs) has shown promising results for speech separation tasks. A limitation of these models is that they have a fixed receptive field (RF). Recent research in speech dereverberation has shown that the optimal RF of a TCN varies with the reverberation characteristics of the speech signal. In this work deformable convolution is proposed as a solution to allow TCN models to have dynamic RFs that can adapt to various reverberation times for reverberant speech separation. The proposed models are capable of achieving an 11.1 dB average scale-invariant signal-to-distortion ratio (SISDR) improvement over the input signal on the WHAMR benchmark. A relatively small deformable TCN model of 1.3M parameters is proposed which gives comparable separation performance to larger and more computationally complex models. William Ravenscroft, Stefan Goetze, Thomas Hain |
ICASSP | 2 |
| 2022 | Non-intrusive Speech Intelligibility Metric Prediction for Hearing Impaired IndividualsabstractThis paper proposes neural models to predict Speech Intelligibility (SI),both by prediction of established SI metrics and of human speech recognition (HSR) on the 1st Clarity Prediction Challenge. Both intrusive and non-intrusive predictors for intrusive SI metrics are trained, then fine tuned on the HSR ground truth. Results are reported on a number of SI metrics, and the model choice for the Clarity challenge submission is explained. Additionally, the relationship between the SI scores in the data and commonly used signal processing metrics which approximate SI are analysed, and some issues emerging from this relationship discussed. It is found that intrusive neural predictors of SI metrics when finetuned on the true HSR scores outperform the non neural challenge baseline. George Close, Samuel Schmück, Stefan Goetze, Thomas Hain |
INTERSPEECH | 3 |
| 2019 | Non-Intrusive Speech Quality Prediction Using Modulation Energies and LSTM-NetworkabstractMany signal processing algorithms have been proposed to improve the quality of speech recorded in the presence of noise and reverberation. Perceptual measures, i.e., listening tests, are usually considered the most reliable way to evaluate the quality of speech processed by such algorithms but are costly and time-consuming. Consequently, speech enhancement algorithms are often evaluated using signal-based measures, which can be either intrusive or non-intrusive. As the computation of intrusive measures requires a reference signal, only non-intrusive measures can be used in applications for which the clean speech signal is not available. However, many existing non-intrusive measures correlate poorly with the perceived speech quality, particularly when applied over a wide range of algorithms or acoustic conditions. In this paper, we propose a novel non-intrusive measure of the quality of processed speech that combines modulation energy features and a recurrent neural network using long short-term memory cells. We collected a dataset of perceptually evaluated signals representing several acoustic conditions and algorithms and used this dataset to train and evaluate the proposed measure. Results show that the proposed measure yields higher correlation with perceptual speech quality than that of benchmark intrusive and non-intrusive measures when considering various categories of algorithms. Although the proposed measure is sensitive to mismatch between training and testing, results show that it is a useful approach to evaluate specific algorithms over a wide range of acoustic conditions and may, thus, become particularly useful for real-time selection of speech enhancement algorithm settings. Benjamin Cauchi, Kai Siedenburg, João Felipe Santos, Tiago H. Falk, Simon Doclo, Stefan Goetze |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2019 | Joint Estimation of Reverberation Time and Early-To-Late Reverberation Ratio From Single-Channel Speech SignalsabstractThe reverberation time (RT) and the early-to-late reverberation ratio (ELR) are two key parameters commonly used to characterize acoustic room environments. In contrast to conventional blind estimation methods that process the two parameters separately, we propose a model for joint estimation to predict the RT and the ELR simultaneously from single-channel speech signals from either full-band or sub-band frequency data, which is referred to as joint room parameter estimator (jROPE). An artificial neural network is employed to learn the mapping from acoustic observations to the RT and the ELR classes. Auditory-inspired acoustic features obtained by temporal modulation filtering of the speech time-frequency representations are used as input for the neural network. Based on an in-depth analysis of the dependency between the RT and the ELR, a two-dimensional (RT, ELR) distribution with constrained boundaries is derived, which is then exploited to evaluate four different configurations for jROPE. Experimental results show that-in comparison to the single-task ROPE system which individually estimates the RT or the ELR-jROPE provides improved results for both tasks in various reverberant and (diffuse) noisy environments. Among the four proposed joint types, the one incorporating multi-task learning with shared input and hidden layers yields the best estimation accuracies on average. When encountering extreme reverberant conditions with RTs and ELRs lying beyond the derived (RT, ELR) distribution, the type considering RT and ELR as a joint parameter performs robustly, in particular. From state-of-the-art algorithms that were tested in the acoustic characterization of environments challenge, jROPE achieves comparable results among the best for all individual tasks (RT and ELR estimation from full-band and sub-band signals). Feifei Xiong, Stefan Goetze, Birger Kollmeier, Bernd T. Meyer |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Exploring Auditory-Inspired Acoustic Features for Room Acoustic Parameter Estimation From Monaural SpeechabstractRoom acoustic parameters that characterize acoustic environments can help to improve signal enhancement algorithms such as for dereverberation, or automatic speech recognition by adapting models to the current parameter set. The reverberation time (RT) and the early-to-late reverberation ratio (ELR) are two key parameters. In this paper, we propose a blind ROom Parameter Estimator (ROPE) based on an artificial neural network that learns the mapping to discrete ranges of the RT and the ELR from single-microphone speech signals. Auditory-inspired acoustic features are used as neural network input, which are generated by a temporal modulation filter bank applied to the speech time-frequency representation. ROPE performance is analyzed in various reverberant environments in both clean and noisy conditions for both fullband and subband RT and ELR estimations. The importance of specific temporal modulation frequencies is analyzed by evaluating the contribution of individual filters to the ROPE performance. Experimental results show that ROPE is robust against different variations caused by room impulse responses (measured versus simulated), mismatched noise levels, and speech variability reflected through different corpora. Compared to state-of-the-art algorithms that were tested in the acoustic characterisation of environments (ACE) challenge, the ROPE model is the only one that is among the best for all individual tasks (RT and ELR estimation from fullband and subband signals). Improved fullband estimations are even obtained by ROPE when integrating speech-related frequency subbands. Furthermore, the model requires the least computational resources with a real time factor that is at least two times faster than competing algorithms. Results are achieved with an average observation window of 3 s, which is important for real-time applications. Feifei Xiong, Stefan Goetze, Birger Kollmeier, Bernd T. Meyer |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Measuring, modelling and predicting perceived reverberationabstractThis paper investigates the relationship between the perceived level of reverberation and parameters measured from the room impulse response (RIR), as well as the design of an instrumental measure that predicts this perceived level. We first present the results of an experimental listening test conducted to assess the level of perceived reverberation in speech captured by a single microphone, before analysing the gathered data to assess the influence of parameters such as the reverberation time (T60) or the direct-to-reverberant ratio (DRR). Secondly, we use the results of this analysis to improve the signal based reverberation decay tail (RDT) measure, previously proposed by the authors to predict the perceived level of reverberation. The accuracy of the proposed measure is evaluated in terms of correlation with the subjective scores and compared to the performance of predictors using parameters extracted from the RIR. Results show that the proposed modifications to the RDT does improve its accuracy. Though still slightly outperformed by measures based on parameters of the RIR, we believe the proposed measure to be useful in scenarios in which the RIR or its parameters are unknown. Hamza A. Javed, Benjamin Cauchi, Simon Doclo, Patrick A. Naylor, Stefan Goetze |
ICASSP | 5 |
| 2017 | Combination strategy based on relative performance monitoring for multi-stream reverberant speech recognitionabstractA multi-stream framework with deep neural network (DNN) classifiers is applied to improve automatic speech recognition (ASR) in environments with different reverberation characteristics. We propose a room parameter estimation model to establish a reliable combination strategy which performs on either DNN posterior probabilities or word lattices. The model is implemented by training a multilayer perceptron incorporating auditory-inspired features in order to distinguish between and generalize to various reverberant conditions, and the model output is shown to be highly correlated to ASR performances between multiple streams, i.e., relative performance monitoring, in contrast to conventional mean temporal distance based performance monitoring for a single stream. Compared to traditional multi-condition training, average relative word error rate improvements of 7.7% and 9.4% have been achieved by the proposed combination strategies performing on posteriors and lattices, respectively, when the multi-stream ASR is tested in known and unknown simulated reverberant environments as well as realistically recorded conditions taken from REVERB Challenge evaluation set. Feifei Xiong, Stefan Goetze, Bernd T. Meyer |
ICASSP | 2 |
| 2017 | On DNN posterior probability combination in multi-stream speech recognition for reverberant environmentsabstractA multi-stream framework with deep neural network (DNN) classifiers has been applied in this paper to improve automatic speech recognition (ASR) performance in environments with different reverberation characteristics. We propose a room parameter estimation model to determine the stream weights for DNN posterior probability combination with the aim of obtaining reliable log-likelihoods for decoding. The model is implemented by training a multi-layer perceptron to distinguish between various reverberant environments. The method is tested in known and unknown environments against approaches based on inverse entropy and autoencoders, with average relative word error rate improvements of 46% and 29%, respectively, when performing multi-stream ASR in different reverberant situations. Feifei Xiong, Stefan Goetze, Bernd T. Meyer |
ICASSP | 2 |
| 2017 | Multi-Channel Speech Enhancement and Amplitude Modulation Analysis for Noise Robust Automatic Speech Recognition
Niko Moritz, Kamil Adiloglu, Jörn Anemüller, Stefan Goetze, Birger Kollmeier |
Comput. Speech Lang. | 4 |
| 2017 | Classifier Architectures for Acoustic Scenes and Events: Implications for DNNs, TDNNs, and Perceptual Features from DCASE 2016abstractThis paper evaluates neural network (NN) based systems and compares them to Gaussian mixture model (GMM) and hidden Markov model (HMM) approaches for acoustic scene classification (SC) and polyphonic acoustic event detection (AED) that are applied to data of the “Detection and Classification of Acoustic Scenes and Events 2016” (DCASE'16) challenge, task 1 and task 3, respectively. For both tasks, the use of deep neural networks (DNNs) and features based on an amplitude modulation filterbank and a Gabor filterbank (GFB) are evaluated and compared to standard approaches. For SC, additionally a time-delay NN approach is proposed that enables analysis of long contextual information similar to recurrent NNs but with training efforts comparable to conventional DNNs. The SC system proposed for task 1 of the DCASE'16 challenge attains a recognition accuracy of 77.5%, which is 5.6% higher compared to the DCASE'16 baseline system. For the AED task, DNNs are adopted in tandem and hybrid approaches, i.e., as part of HMM-based systems. These systems are evaluated for the polyphonic data of task 3 from the DCASE'16 challenge. Several strategies to address the issue of polyphony are considered. It is shown that DNN-based systems perform less accurate than the traditional systems for this task. Best results are achieved using GFB features in combination with a multiclass GMM-HMM back end. Jens Schröder, Niko Moritz, Jörn Anemüller, Stefan Goetze, Birger Kollmeier |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2016 | Perceptual and instrumental evaluation of the perceived level of reverberationabstractPerceptual measures are usually considered more reliable than instrumental measures for evaluating the perceived level of reverberation. However, such measures are costly in both time and money, and, due to variations in stimuli or assessors, the resulting data is not always statistically significant. Therefore, an efficient perceptual measure of the perceived level of reverberation is needed. We compare the use of a multiple stimuli test with the use of pairwise comparison for the evaluation of the perceived level of reverberation. The results suggest that using multiple stimuli is preferable to pairwise comparison as long as the number of conditions to be compared is not too large. Additionally, we use the results from the conducted perceptual measurements to examine the reliability of existing instrumental measures of the perceived level of reverberation. Our observations show which instrumental measures are effective in highlighting differences between RIR characteristics and which ones have to be preferred if one aims at predicting the level of reverberation perceived by a human assessor. Benjamin Cauchi, Hamza A. Javed, Timo Gerkmann, Simon Doclo, Stefan Goetze, Patrick A. Naylor |
ICASSP | 5 |
| 2016 | Classification of human cough signals using spectro-temporal Gabor filterbank featuresabstractThis contribution investigates the use of features derived from a Gabor filterbank (GFB) for the application of acoustic cough classification. Gabor filters are two-dimensional filters that decompose the spectro-temporal power density further into components which capture spectral, temporal and joint spectro-temporal modulation patterns. The proposed GFB feature extraction scheme in combination with Gaussian mixture model (GMM) and hidden Markov model (HMM) classifier back-ends is evaluated using a cough database recorded by a phone hotline. The database is composed of two kind of coughs, i.e., dry and productive cough, and other sounds, e.g. speech. Based on these data, we show that GFB features result in better recognition performance than the common Mel-frequency cepstral coefficient (MFCC) baseline for the given task of cough classification. Furthermore, results indicate that GMMs are preferable to HMMs for this kind of data. Jens Schröder, Jörn Anemüller, Stefan Goetze |
ICASSP | 3 |
| 2015 | A CHiME-3 challenge system: Long-term acoustic features for noise robust automatic speech recognitionabstractThe paper describes an automatic speech recognition (ASR) system for the 3rd CHiME challenge that addresses noisy acoustic scenes within public environments. The proposed system includes a multi-channel speech enhancement front-end including a microphone channel failure detection method that is based on cross-comparing the modulation spectra of speech to detect erroneous microphone recordings. The main focus of the submission is the investigation of the amplitude modulation filter bank (AMFB) as a method to extract long-term acoustic cues prior to a Gaussian mixture model (GMM) or deep neural network (DNN) based ASR classifier. It is shown that AMFB features outperform the commonly used frame splicing technique of filter bank features even on a performance optimized ASR challenge system. I.e., temporal analysis of speech by hand-crafted and auditory motivated AMFBs is shown to be more robust compared to a data-driven method based on extracting temporal dynamics with a DNN. Our final ASR system, which additionally includes adaptation of acoustic features to speaker characteristics, achieves an absolute word error rate reduction of approx. 21.53 % relative to the best CHiME-3 baseline system on the "real" test condition. Niko Moritz, Stephan Gerlach, Kamil Adiloglu, Jörn Anemüller, Birger Kollmeier, Stefan Goetze |
ASRU | 6 |
| 2015 | A study on joint beamforming and spectral enhancement for robust speech recognition in reverberant environmentsabstractThis work evaluates multi-microphone beamforming and single-microphone spectral enhancement strategies to alleviate the reverberation effect for robust automatic speech recognition (ASR) systems in different reverberant environments characterized by different reverberation times T60 and direct-to-reverberation ratios (DRRs). The systems consist of minimum variance distortionless response (MVDR) beamformers in combination with minimum mean square error (MMSE) estimators, and late reverberation spectral variance (LRSV) estimators, the latter employing a generalized model of the room impulse response (RIR). Various system architectures are analyzed with a focus on optimal speech recognition performance. The system combining an MVDR beamformer and a subsequent MMSE estimator was found to lead to the best results, with relative reductions of 27.7% compared to the baseline system. This is attributed to a more accurate LRSV estimate from spatial averaging and diffuse field refinement for the MMSE estimator. Feifei Xiong, Bernd T. Meyer, Stefan Goetze |
ICASSP | 3 |
| 2015 | Reduction of Gaussian, Supergaussian, and Impulsive Noise by Interpolation of the Binary Mask ResidualabstractIn this paper, we present a new approach for noise reduction. A binary time-frequency (T-F) masking threshold criterion is proposed and analyzed with respect to the average spectra of music and noise disturbances. Modified autoregressive (AR) detection and AR interpolation are then applied to the residual signal of the binary masking process. The proposed method is able to reduce supergaussian and impulsive noise while ensuring preservation of the desired signal, which is crucial for professional high-quality audio restoration, and it is also suitable for Gaussian noise to a certain extent. The approach is compared to a state-of-the-art restoration algorithm by means of the objective measures signal-to-noise ratio (SNR) improvement and perceptual quality, and by subjective listening tests. The objective results as well as the listening tests show that the proposed algorithm is especially suited for supergaussian, grainy-sounding noise types, e.g., optical soundtrack noise of celluloid movie footage, or rain noise. Marco Ruhland, Jörg Bitzer, Matthias Brandt, Stefan Goetze |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | Spectro-Temporal Gabor Filterbank Features for Acoustic Event DetectionabstractAlgorithms for the automatic detection and recognition of acoustic events are increasingly gaining relevance for the reliable and robust functioning of consumer, assistive and monitoring systems. The extraction of appropriate task relevant acoustic features from the raw sound signal clearly influences performance of subsequent statistical classification, in particular in adverse acoustic situations. The present contribution investigates the use of biologically-inspired features, derived from a filterbank of two-dimensional Gabor functions, that decompose the spectro-temporal power density into components which capture spectral, temporal and joint spectro-temporal modulation patterns. It is hypothesized that the comparably large joint spectral and temporal extent of these Gabor functions results in features that allow for robust classification. Evaluation of the proposed feature extraction scheme together with an hidden Markov model (HMM) classifier is conducted on two corpora comprising acoustic events in realistic adverse conditions from the D-CASE and CLEAR'07 evaluation campaigns. Relevance of each Gabor filter for classification is analyzed and an optimized parameter set for the Gabor filterbank (GFB) is identified. Performance of the optimized GFB is evaluated in comparison to other state-of-the-art algorithms on isolated event classification and on the full acoustic event detection (AED) including joint classification and temporal segmentation of events. Results show that Gabor features result in a signal representation that exhibits separated average class-specific patterns. An improvement in classification accuracy of up to 26% relative to the Mel-frequency cepstral coefficient (MFCC) baseline is obtained with the optimized GFB. Further experiments demonstrate that this improvement cannot be explained by purely temporal or purely spectral Gabor basis functions. Rather, a GFB with features extending in joint spectro-temporal directions is required to obtain optimum performance. Performance on AED with the D-CASE challenge dataset is shown to improve on previous algorithms from the recent literature. Jens Schröder, Stefan Goetze, Jörn Anemüller |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Estimating room acoustic parameters for speech recognizer adaptation and combination in reverberant environmentsabstractThis work analyzes the influence of reverberation on automatic speech recognition (ASR) systems and how to compensate its influence, with special focus on the important acoustical parameters i.e. room reverberation time T60and clarity index C50. A multilayer perceptron (MLP) using features of a spectro-temporal filter bank as input is employed to identify the acoustic conditions spanning various reverberant scenarios. The posterior probabilities of the MLP are used to design a novel selection scheme for adaptation in a cluster-based manner and for system combination achieved by recognizer output voting error reduction (ROVER). A comparison of word error rates is performed considering different training modes, and an average relative improvement of 7.1% is obtained by the proposed system compared to conventional multistyle training. Feifei Xiong, Stefan Goetze, Bernd T. Meyer |
ICASSP | 2 |
| 2013 | A perceptually constrained channel shortening technique for speech dereverberationabstractThe objective of acoustic multichannel equalization is to design a reshaping filter that reduces reverberation, improves the perceptual speech quality, and is robust to errors in the estimated room impulse responses (RIRs). Although the channel shortening (CS) technique has been shown to be effective in achieving dereverberation, it may fail to preserve the natural shape of an RIR leading to speech quality degradation. Furthermore, CS yields multiple reshaping filters that satisfy its optimization criterion but result in a different perceptual speech quality. In this paper, we propose a robust perceptually constrained channel shortening technique (PeCCS) that resolves the selection ambiguity of CS and leads to joint dereverberation and speech quality preservation. Simulation results for erroneously estimated RIRs show that PeCCS preserves the perceptual speech quality and results in a higher reverberant tail suppression than other state-of-the-art techniques, such as CS and the regularized partial multichannel equalization technique based on the multiple-input/output inverse theorem (P-MINT). Ina Kodrasi, Stefan Goetze, Simon Doclo |
ICASSP | 2 |
| 2013 | Automatic acoustic siren detection in traffic noise by part-based modelsabstractState-of-the-art classifiers like hidden Markov models (HMMs) in combination with mel-frequency cepstral coefficients (MFCCs) are flexible in time but rigid in the spectral dimension. In contrast, part-based models (PBMs) originally proposed in computer vision consist of parts in a fully deformable configuration. The present contribution proposes to employ PBMs in the spectro-temporal domain for detection of emergency siren sounds in traffic noise,standard generative training resulting in a classifier that is robust to shifts in frequency induced, e.g., by Doppler-shift effects. Two improvements over standard machine learning techniques for PBM estimation are proposed: (i) Spectro-temporal part (“appearance”) extraction is initialized by interest point detection instead of random initialization and (ii) a discriminative training approach in addition to standard generative training is implemented. Evaluation with self-recorded police sirens and traffic noise gathered on-line demonstrates that PBMs are successful in acoustic siren detection. One hand-labeled and two machine learned PBMs are compared to standard HMMs employing mel-spectrograms and MFCCs in clean and multi condition (multiple SNR) training settings. Results show that in clean condition training, hand-labeled PBMs and HMMs outperform machine-learned PBMs already for test data with moderate additive noise. In multi condition training, the machine learned PBMs outperform HMMs on most SNRs, achieving high accuracies and being nearly optimal up to 5 dB SNR. Thus, our simulation results show that PBMs are a promising approach for acoustic event detection (AED). Jens Schröder, Stefan Goetze, Volker Grutzmacher, Jörn Anemüller |
ICASSP | 2 |
| 2013 | Blind estimation of reverberation time based on spectro-temporal modulation filteringabstractA novel method for blind estimation of the reverberation time (RT60) is proposed based on applying spectro-temporal modulation filters to time-frequency representations. 2D-Gabor filters arranged in a filterbank enable an analysis of the properties of temporal, spectral, and spectro-temporal filtering for this task. Features are used as input to a multi-layer perceptron (MLP) classifier combined with a simple decision rule that attributes a specific RT60 to a given utterance and allows to assess the reliability of the approach for different resolutions of RT60 classification. While the filter set including temporal, spectral, and spectro-temporal filters already outperforms an MFCC baseline, the error rates are further reduced when relying on diagonal spectro-temporal filters alone. The average error rate is 1.9% for the best feature set, which corresponds to a relative reduction of 58.3% compared to the MFCC baseline for RT60s in 0.1 s resolution. Feifei Xiong, Stefan Goetze, Bernd T. Meyer |
ICASSP | 2 |
| 2013 | Regularization for Partial Multichannel Equalization for Speech DereverberationabstractAcoustic multichannel equalization techniques such as the multiple-input/output inverse theorem (MINT), which aim to equalize the room impulse responses (RIRs) between the source and the microphone array, are known to be highly sensitive to RIR estimation errors. To increase robustness, it has been proposed to incorporate regularization in order to decrease the energy of the equalization filters. In addition, more robust partial multichannel equalization techniques such as relaxed multichannel least-squares (RMCLS) and channel shortening (CS) have recently been proposed. In this paper, we propose a partial multichannel equalization technique based on MINT (P-MINT) which aims to shorten the RIR. Furthermore, we investigate the effectiveness of incorporating regularization to further increase the robustness of P-MINT and the aforementioned partial multichannel equalization techniques, i.e., RMCLS and CS. In addition, we introduce an automatic non-intrusive procedure for determining the regularization parameter based on the L-curve. Simulation results using measured RIRs show that incorporating regularization in P-MINT yields a significant performance improvement in the presence of RIR estimation errors, whereas a smaller performance improvement is observed when incorporating regularization in RMCLS and CS. Furthermore, it is shown that the intrusively regularized P-MINT technique outperforms all other investigated intrusively regularized multichannel equalization techniques in terms of perceptual speech quality (PESQ). Finally, it is shown that the automatic non-intrusive regularization parameter in regularized P-MINT leads to a very similar performance as the intrusively determined optimal regularization parameter, making regularized P-MINT a robust, perceptually advantageous, and practically applicable multichannel equalization technique for speech dereverberation. Ina Kodrasi, Stefan Goetze, Simon Doclo |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | System identification for listening-room compensation by means of acoustic echo cancellation and acoustic echo suppression filtersabstractSubsystems for dereverberation and acoustic echo cancellation (AEC)/acoustic echo suppression (AES) are important components in high-quality hands-free telecommunication systems. This contribution describes and analyzes a combined system for dereverberation and AEC/AES. The system identification inherently achieved by the AEC/AES system is used for the design of the room impulse response (RIR) equalization filter, i.e. the listening-room compensation (LRC) system. We use complex RIR smoothing and decoupled filtered-X least-mean-squares (dFxLMS) gradient algorithm for LRC and a combined AEC/AES system for the system identification necessary for the LRC filter design. The performance of the combined system and the mutual influences of LRC and AEC/AES are analyzed. Feifei Xiong, Jens-E. Appell, Stefan Goetze |
ICASSP | 3 |
| 2010 | Quality assessment for listening-room compensation algorithmsabstractIn this contribution various objective measures that can be used to evaluate speech dereverberation algorithms by means of listening-room compensation (LRC) are compared to subjective listening tests. It is shown that technical measures describing the impulse responses are suitable for evaluation of such algorithms. Most signal-based objective measures fail to judge the specific distortions that may be introduced by LRC algorithms like late reverberation since these artifacts are small in amplitude but perceptually relevant due to the loss of masking of the room impulse response. Only one signal-based measure, the so-called perceptual similarity measure (PSM), showed high correlation with subjective rating for the given test setup. Stefan Goetze, Eugen Albertin, Markus Kallinger, Alfred Mertins, Karl-Dirk Kammeyer |
ICASSP | 1 |
| 2010 | Automatic Live Monitoring of Communication Quality for Normal-Hearing and Hearing-Impaired Listeners
Jan Rennies, Eugen Albertin, Stefan Goetze, Jens-E. Appell |
ICCHP (2) | 3 |
| 2008 | Objective perceptual quality assessment for self-steering binaural hearing aid microphone arraysabstractIn this study a self-steering beamformer with binaural output for a head-worn microphone array is investigated in simulated and real- world conditions. The influence of the underlying sound propagation model on the estimation accuracy of the direction of arrival (DOA) estimation algorithm and the overall performance of the combined DOA-beamformer-system is evaluated. For this, technical performance measures as well as objective quality measures based on perceptual models of the auditory system are used. The self-steering beamformer showed better performance than a beamformer with fixed look-direction for SNR values above -2 dB if the propagation model includes at least a coarse head model. Thomas Rohdenburg, Stefan Goetze, Volker Hohmann, Karl-Dirk Kammeyer, Birger Kollmeier |
ICASSP | 2 |
| 2007 | Optimization of Gabor Features for Text-Independent Speaker IdentificationabstractFor text-independent speaker identification a prominent combination is to use Gaussian mixture models (GMM) for classification while relying on Mel-frequency cepstral coefficients (MFCC) as features. To take temporal information into account the time difference of features of adjacent speech frames are appended to the initial features. In this paper we investigate the applicability of spectro-temporal features obtained from Gabor-filters and present an algorithm for optimizing the possible parameters. Simulation results on a database show that spectro-temporal features achieve higher recognition rates than purely temporal features for clean speech as well as for disturbed speech. Volker Mildner, Stefan Goetze, Karl-Dirk Kammeyer, Alfred Mertins |
ISCAS | 2 |