Ina Kodrasi

dblp:26/9876 · DBLP profile ↗
← Back
34ranked-venue papers
18as first author
12since 2021 · last 2025
0000-0002-1747-1322ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 13 first-author · 11 since 2021Artificial intelligence and machine learning · 13 · 7 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Multiview Canonical Correlation Analysis for Automatic Pathological Speech Detection
abstract
Recently proposed automatic pathological speech detection approaches rely on spectrogram input representations or wav2vec2 embeddings. These representations may contain pathology-irrelevant uncorrelated information, such as changing phonetic content or variations in speaking style across time, which can adversely affect classification performance. To address this issue, we propose to use Multiview Canonical Correlation Analysis (MCCA) on these input representations prior to automatic pathological speech detection. Our results demonstrate that unlike other dimensionality reduction techniques, the use of MCCA leads to a considerable improvement in pathological speech detection performance by eliminating uncorrelated information present in the input representations. Employing MCCA with traditional classifiers yields a comparable or higher performance than using sophisticated architectures, while preserving the representation structure and providing interpretability.
Yacouba Kaloga, Shakeel A. Sheikh, Ina Kodrasi
ICASSP3
2025 Graph Neural Networks for Parkinson's Disease Detection
abstract
Despite the promising performance of state-of-the-art approaches for Parkinson’s Disease (PD) detection, these approaches often analyze individual speech segments in isolation, which can lead to sub-optimal results. Dysarthric cues that characterize speech impairments from PD patients are expected to be related across segments from different speakers. Isolated segment analysis fails to exploit these inter-segment relationships. Additionally, not all speech segments from PD patients exhibit clear dysarthric symptoms, introducing label noise that can negatively affect the performance and generalizability of current approaches. To address these challenges, we propose a novel PD detection framework utilizing Graph Convolutional Networks (GCNs). By representing speech segments as nodes and capturing the similarity between segments through edges, our GCN model facilitates the aggregation of dysarthric cues across the graph, effectively exploiting segment relationships and mitigating the impact of label noise. Experimental results demonstrate the advantages of the proposed GCN model for PD detection and provide insights into its underlying mechanisms.
Shakeel A. Sheikh, Yacouba Kaloga, Md. Sahidullah, Ina Kodrasi
ICASSP4
2025 Latent Space Factorization in LoRA
abstract
Low-rank adaptation (LoRA) is a widely used method for parameter-efficient finetuning. However, existing LoRA variants lack mechanisms to explicitly disambiguate task-relevant information within the learned low-rank subspace, potentially limiting downstream performance. We propose Factorized Variational Autoencoder LoRA (FVAE-LoRA), which leverages a VAE to learn two distinct latent spaces. Our novel Evidence Lower Bound formulation explicitly promotes factorization between the latent spaces, dedicating one latent space to task-salient features and the other to residual information. Extensive experiments on text, audio, and image tasks demonstrate that FVAE-LoRA consistently outperforms standard LoRA. Moreover, spurious correlation evaluations confirm that FVAE-LoRA better isolates task-relevant signals, leading to improved robustness under distribution shifts. Our code is publicly available at: https://github.com/idiap/FVAE-LoRA
Shashi Kumar, Yacouba Kaloga, John Mitros, Petr Motlícek, Ina Kodrasi
NeurIPS5
2024 Adversarial Robustness Analysis in Automatic Pathological Speech Detection Approaches
Mahdi Amiri, Ina Kodrasi
INTERSPEECH2
2023 On Using the UA-Speech and Torgo Databases to Validate Automatic Dysarthric Speech Classification Approaches
abstract
Although the UA-Speech and TORGO databases of control and dysarthric speech are invaluable resources made available to the research community with the objective of developing robust automatic speech recognition systems, they have also been used to validate a considerable number of automatic dysarthric speech classification approaches. Such approaches typically rely on the underlying assumption that recordings from control and dysarthric speakers are collected in the same noiseless environment using the same recording setup. In this paper, we show that this assumption is violated for the UA-Speech and TORGO databases. Using voice activity detection to extract speech and non-speech segments, we show that the majority of state-of-the-art dysarthria classification approaches achieve the same or a considerably better performance when using the non-speech segments of these databases than when using the speech segments. These results demonstrate that such approaches trained and validated on the UA-Speech and TORGO databases are potentially learning characteristics of the recording environment or setup rather than dysarthric speech characteristics. We hope that these results raise awareness in the research community about the importance of the quality of recordings when developing and evaluating automatic dysarthria classification approaches.
Guilherme Schu, Parvaneh Janbakhshi, Ina Kodrasi
ICASSP3
2022 Experimental Investigation on STFT Phase Representations for Deep Learning-Based Dysarthric Speech Detection
abstract
Mainstream deep learning-based dysarthric speech detection approaches typically rely on processing the magnitude spectrum of the short-time Fourier transform of input signals, while ignoring the phase spectrum. Although considerable insight about the structure of a signal can be obtained from the magnitude spectrum, the phase spectrum also contains inherent structures which are not immediately apparent due to phase discontinuity. To reveal meaningful phase structures, alternative phase representations such as the modified group delay (MGD) and instantaneous frequency (IF) spectra have been investigated in several applications. The objective of this paper is to investigate the applicability of the unprocessed phase, MGD, and IF spectra for dysarthric speech detection. Experimental results show that dysarthric cues are present in all considered phase representations. Further, it is shown that using phase representations as complementary features to the magnitude spectrum is beneficial for deep learning-based dysarthric speech detection, with the combination of magnitude and IF spectra yielding a high performance. The presented results should raise awareness in the research community about the potential of the phase spectrum for dysarthric speech detection and motivate research into novel architectures which optimally exploit magnitude and phase information.
Parvaneh Janbakhshi, Ina Kodrasi
ICASSP2
2022 Comparison of 5 methods for the evaluation of intelligibility in mild to moderate French dysarthric speech
abstract
Altered quality of the phonetic-acoustic information in the speech signal in the case of motor speech disorders may reduce its intelligibility. Monitoring intelligibility is part of the standard clinical assessment of patients. It is also a valuable tool to index the evolution of the speech disorder. However, measuring intelligibility raises methodological debates concerning: the type of linguistic material on which the assessment is based (non-words, words, continuous speech), the evaluation protocol and type of scores (scale-based rating, transcription or recognition tests), and the advantages and disadvantages of listener vs. automatic-based approaches (subjective vs. objective, expertise level, types of models used). In this paper, the intelligibility of the speech of 32 French patients presenting mild to moderate dysarthria and 17 elderly speakers is assessed with five different methods: impressionistic clinician judgment on continuous speech, number of words recognized in an interactive face-to-face setting and in an on-line testing of the same material by 75 judges, automatic feature-based and automatic speech recognition-based methods (both on short sentences). The implications of the different methods for clinical practice are discussed.
Cécile Fougeron, Nicolas Audibert, Ina Kodrasi, Parvaneh Janbakhshi, Michaela Pernon, Nathalie Lévêque, Stephanie Borel, Marina Laganaro, Hervé Bourlard, Frédéric Assal
INTERSPEECH3
2022 Adversarial-Free Speaker Identity-Invariant Representation Learning for Automatic Dysarthric Speech Classification
abstract
Speech representations which are robust to pathology-unrelated cues such as speaker identity information have been shown to be advantageous for automatic dysarthric speech classification. A recently proposed technique to learn speaker identity-invariant representations for dysarthric speech classification is based on adversarial training. However, adversarial training can be challenging, unstable, and sensitive to training parameters. To avoid adversarial training, in this paper we propose to learn speaker-identity invariant representations exploiting a feature separation framework relying on mutual information minimization. Experimental results on a database of neurotypical and dysarthric speech show that the proposed adversarial-free framework successfully learns speaker identity-invariant representations. Further, it is shown that such representations result in a similar dysarthric speech classification performance as the representations obtained using adversarial training, while the training procedure is more stable and less sensitive to training parameters.
Parvaneh Janbakhshi, Ina Kodrasi
INTERSPEECH2
2021 Automatic Dysarthric Speech Detection Exploiting Pairwise Distance-Based Convolutional Neural Networks
abstract
Automatic dysarthric speech detection can provide reliable and cost-effective computer-aided tools to assist the clinical diagnosis and management of dysarthria. In this paper we propose a novel automatic dysarthric speech detection approach based on analyses of pairwise distance matrices using convolutional neural networks (CNNs). We represent utterances through articulatory posteriors and consider pairs of phonetically-balanced representations, with one representation from a healthy speaker (i.e., the reference representation) and the other representation from the test speaker (i.e., test representation). Given such pairs of reference and test representations, features are first extracted using a feature extraction front-end, a frame-level distance matrix is computed, and the obtained distance matrix is considered as an image by a CNN-based binary classifier. The feature extraction, distance matrix computation, and CNN-based classifier are jointly optimized in an end-to-end framework. Experimental results on two databases of healthy and dysarthric speakers for different languages and pathologies show that the proposed approach yields a high dysarthric speech detection performance, outperforming other CNN-based baseline approaches.
Parvaneh Janbakhshi, Ina Kodrasi, Hervé Bourlard
ICASSP2
2021 Automatic And Perceptual Discrimination Between Dysarthria, Apraxia of Speech, and Neurotypical Speech
abstract
Automatic techniques in the context of motor speech disorders (MSDs) are typically two-class techniques aiming to discriminate between dysarthria and neurotypical speech or between dysarthria and apraxia of speech (AoS). Further, although such techniques are proposed to support the perceptual assessment of clinicians, the automatic and perceptual classification accuracy has never been compared. In this paper, we investigate a three-class automatic technique and a set of handcrafted features for the discrimination of dysarthria, AoS and neurotypical speech. Instead of following the commonly used One-versus-One or One-versus-Rest approaches for multi-class classification, a hierarchical approach is proposed. Further, a perceptual study is conducted where speech and language pathologists are asked to listen to recordings of dysarthria, AoS, and neurotypical speech and decide which class the recordings belong to. The proposed automatic technique is evaluated on the same recordings and the automatic and perceptual classification performance are compared. The presented results show that the hierarchical classification approach yields a higher classification accuracy than baseline One-versus-One and One-versus-Rest approaches. Further, the presented results show that the automatic approach yields a higher classification accuracy than the perceptual assessment of speech and language pathologists, demonstrating the potential advantages of integrating automatic tools in clinical practice.
Ina Kodrasi, Michaela Pernon, Marina Laganaro, Hervé Bourlard
ICASSP1
2021 Subspace-Based Learning for Automatic Dysarthric Speech Detection
abstract
To assist the clinical diagnosis and treatment of speech dysarthria, automatic dysarthric speech detection techniques providing reliable and cost-effective assessment are indispensable. Based on clinical evidence on spectro-temporal distortions associated with dysarthric speech, we propose to automatically discriminate between healthy and dysarthric speakers exploiting spectro-temporal subspaces of speech. Spectro-temporal subspaces are extracted using singular value decomposition, and dysarthric speech detection is achieved by applying a subspace-based discriminant analysis. Experimental results on databases of healthy and dysarthric speakers for different languages and pathologies show that the proposed subspace-based approach using temporal subspaces is more advantageous than using spectral subspaces, also outperforming several state-of-the-art automatic dysarthric speech detection techniques.
Parvaneh Janbakhshi, Ina Kodrasi, Hervé Bourlard
IEEE Signal Process. Lett.2
2021 Temporal Envelope and Fine Structure Cues for Dysarthric Speech Detection Using CNNs
abstract
Deep learning-based techniques for automatic dysarthric speech detection have recently attracted interest in the research community. State-of-the-art techniques typically learn neurotypical and dysarthric discriminative representations by processing time-frequency input representations such as the magnitude spectrum of the short-time Fourier transform (STFT). Although these techniques are expected to leverage perceptual dysarthric cues, representations such as the magnitude spectrum of the STFT do not necessarily convey perceptual aspects of complex sounds. Inspired by the temporal processing mechanisms of the human auditory system, in this paper we factor signals into the product of a slowly varying envelope and a rapidly varying fine structure. Separately exploiting the different perceptual cues present in the envelope (i.e., phonetic information, stress, and voicing) and fine structure (i.e., pitch, vowel quality, and breathiness), two discriminative representations are learned through a convolutional neural network and used for automatic dysarthric speech detection. Experimental results show that processing both the envelope and fine structure representations yields a considerably better dysarthric speech detection performance than processing only the envelope, fine structure, or magnitude spectrum of the STFT representation.
Ina Kodrasi
IEEE Signal Process. Lett.1
2020 Synthetic Speech References for Automatic Pathological Speech Intelligibility Assessment
abstract
Automatic pathological speech intelligibility measures are crucial to assist the clinical diagnosis and treatment of speech disorders. The recently proposed pathological short-time objective intelligibility (P-ESTOI) measure was shown to be very advantageous, yielding a high performance for several speech pathologies. However, to assess the intelligibility of an utterance from a patient, P-ESTOI relies on the availability of recordings of the same utterance by several healthy speakers such that an intelligible reference model can be created. Such recordings are not always easily available, limiting the practical applicability of P-ESTOI. To be able to use P-ESTOI in such scenarios, in this paper we propose to use synthetic speech generated by state-of-the-art high-quality text-to-speech systems to create an intelligible reference model. Experimental results on a database of Cerebral Palsy patients show that the performance of P-ESTOI using synthetic speech references is comparable to using natural speech references, making P-ESTOI a flexible measure which does not require healthy speech recordings and which outperforms state-of-the-art pathological speech intelligibility measures.
Parvaneh Janbakhshi, Ina Kodrasi, Hervé Bourlard
ICASSP2
2020 Automatic Discrimination of Apraxia of Speech and Dysarthria Using a Minimalistic Set of Handcrafted Features
abstract
To assist clinicians in the differential diagnosis and treatment of motor speech disorders, it is imperative to establish objective tools which can reliably characterize different subtypes of disorders such as apraxia of speech (AoS) and dysarthria.Objective tools in the context of speech disorders typically rely on thousands of acoustic features, which raises the risk of difficulties in the interpretation of the underlying mechanisms, overadaptation to training data, and weak generalization capabilities to test data.Seeking to use a small number of acoustic features and motivated by the clinical-perceptual signs used for the differential diagnosis of AoS and dysarthria, we propose to characterize differences between AoS and dysarthria using only six handcrafted acoustic features, with three features reflecting segmental distortions, two features reflecting loudness and hypernasality, and one feature reflecting syllabification.These three different sets of features are used to separately train three classifiers.At test time, the decisions of the three classifiers are combined through a simple majority voting scheme.Preliminary results show that the proposed approach achieves a discrimination accuracy of 90%, outperforming using state-of-the-art features such as openSMILE which yield a discrimination accuracy of 65%.
Ina Kodrasi, Michaela Pernon, Marina Laganaro, Hervé Bourlard
INTERSPEECH1
2020 Automatic Pathological Speech Intelligibility Assessment Exploiting Subspace-Based Analyses
abstract
Competitive state-of-the-art automatic pathological speech intelligibility measures typically rely on regression training on a large number of features, require a large amount of healthy speech training data, or are applicable only to phonetically balanced scenarios where healthy and pathological speakers utter the same utterances. As a result, their performance in unseen data is unsatisfactory, and they cannot be used in low-resource languages or in phonetically unbalanced scenarios. To overcome these drawbacks, we propose a subspace-based intelligibility (SBI) measure. The SBI measure operates based on the hypothesis that dominant spectral patterns of pathological speech differ from intelligible speech (where the pathological and intelligible speech signals do not need to match in phonetic content), with the difference increasing as pathological speech intelligibility decreases. The SBI measure uses a minimal number of speech recordings to compute dominant spectral basis vectors spanning intelligible and pathological speech. The subspaces spanned by the intelligible and pathological spectral basis vectors are compared to each other through a subspace distance measure, which is directly used (i.e., without any training) as the pathological speech intelligibility estimate. Exploiting psychoacoustic evidence on the importance of spectral modulation cues to the perceived speech intelligibility and clinical evidence on the degradation of these cues in pathological speech, we show that the power of the proposed SBI measure lies in capturing the effect of spectral modulation degradation. To be able to additionally track possible degradations in the temporal structure of the pathological speech signal, we also propose two extensions of the SBI measure by incorporating short-time temporal information. Experimental results for different languages and speech pathologies show that the proposed intelligibility measures yield high and significant correlations with subjective intelligibility ratings, while not requiring any regression training or a large number of healthy speech recordings and being applicable to phonetically unbalanced scenarios.
Parvaneh Janbakhshi, Ina Kodrasi, Hervé Bourlard
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Spectro-Temporal Sparsity Characterization for Dysarthric Speech Detection
abstract
To assist the clinical diagnosis and treatment of neurological diseases that cause speech dysarthria such as Parkinson's disease (PD), it is of paramount importance to craft robust features which can be used to automatically discriminate between healthy and dysarthric speech. Since dysarthric speech of patients suffering from PD is breathy, semi-whispery, and is characterized by abnormal pauses and imprecise articulation, it can be expected that its spectro-temporal sparsity differs from the spectro-temporal sparsity of healthy speech. While we have recently successfully used temporal sparsity characterization for dysarthric speech detection, characterizing spectral sparsity poses the challenge of constructing a valid feature vector from signals with a different number of unaligned time frames. Further, although several non-parametric and parametric measures of sparsity exist, it is unknown which sparsity measure yields the best performance in the context of dysarthric speech detection. The objective of this paper is to demonstrate the advantages of spectro-temporal sparsity characterization for automatic dysarthric speech detection. To this end, we first provide a numerical analysis of the suitability of different non-parametric and parametric measures (i.e., l1-norm, kurtosis, Shannon entropy, Gini index, shape parameter of a Chi distribution, and shape parameter of a Weibull distribution) for sparsity characterization. It is shown that kurtosis, the Gini index, and the parametric sparsity measures are advantageous sparsity measures, whereas the l1-norm and entropy measures fail to robustly characterize the temporal sparsity of signals with a different number of time frames. Second, we propose to characterize the spectral sparsity of an utterance by initially time-aligning it to the same utterance uttered by a (arbitrarily selected) reference speaker using dynamic time warping. Experimental results on a Spanish database of healthy and dysarthric speech show that estimating the spectro-temporal sparsity using the Gini index or the parametric sparsity measures and using it as a feature in a support vector machine results in a high classification accuracy of 83.3%.
Ina Kodrasi, Hervé Bourlard
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Pathological Speech Intelligibility Assessment Based on the Short-time Objective Intelligibility Measure
abstract
Impaired speech intelligibility in motor speech disorders arising due to neurological diseases negatively affects the communication ability and quality of life of patients. Reliable and cost-effective measures to automatically assess speech intelligibility are necessary for the management of such disorders. In this paper, we propose to automatically assess the intelligibility of pathological speech based on short-time objective intelligibility measures typically used in speech enhancement, which however require a reference signal that is time-aligned to the test signal. We propose a method to create an utterance-dependent reference signal of intelligible speech from multiple healthy speakers. In order to assess intelligibility, the pathological speech signal is aligned to the created reference signal using dynamic time warping and the divergence between the two signals is quantified using either the short-time or the spectral correlation. Experiments on databases of English and French patients suffering from Cerebral Palsy and Amyotrophic Lateral Sclerosis show that the proposed intelligibility measures can obtain a high correlation with subjective intelligibility ratings, outperforming several state-of-the-art pathological speech intelligibility measures.
Parvaneh Janbakhshi, Ina Kodrasi, Hervé Bourlard
ICASSP2
2019 Super-gaussianity of Speech Spectral Coefficients as a Potential Biomarker for Dysarthric Speech Detection
abstract
Parkinson's disease (PD) and Amyotrophic Lateral Sclerosis (ALS) are progressive neurodegenerative diseases which, among other symptoms, cause dysarthria of speech. To assist the clinical diagnosis and treatment of neurological diseases, several studies have addressed the characterization and classification of healthy and dysarthric speech. However, most contributions deal with PD speech, with significantly fewer results presented for ALS speech. The objective of this paper is to show that ALS speech has a similar statistical distribution as PD speech, with the complex spectral coefficients being significantly less super-Gaussian than healthy speech spectral coefficients. In addition, a method to exploit the super-Gaussianity of speech signals as a feature to classify healthy and dysarthric speech is presented and evaluated. The proposed approach is evaluated on a French database of healthy and dysarthric (PD and ALS) speech. Experimental results show that the use of the super-Gaussianity of speech signals yields a significantly higher classification accuracy than state-of-the-art features such as fundamental frequency, jitter, shimmer, harmonics-to-noise ratio, or Mel frequency cepstral coefficients.
Ina Kodrasi, Hervé Bourlard
ICASSP1
2019 Joint Estimation of RETF Vector and Power Spectral Densities for Speech Enhancement Based on Alternating Least Squares
abstract
The multi-channel Wiener filter (MWF) is a well-known multi-microphone speech enhancement technique, aiming at improving the quality of the recorded speech signals in noisy and reverberant environments. Assuming that reverberation and ambient noise can be modeled as a diffuse sound field and the spatial coherence of the residual noise is known, the MWF requires estimates of the relative early transfer function (RETF) vector of the target speaker as well as the power spectral densities (PSDs) of the target, diffuse and residual noise component. RETF vector and PSD estimation is often decoupled, where one quantity is estimated independently of the other quantity. In this paper, we propose to jointly estimate the RETF vector and all PSDs by minimizing the Frobenius norm of a model-based error matrix using an alternating least squares method. Experimental results using different dynamic acoustic scenarios with a moving speaker show that the proposed method leads to a larger MWF performance than a state-of-the-art method based on covariance whitening.
Marvin Tammen, Simon Doclo, Ina Kodrasi
ICASSP3
2019 Spectral Subspace Analysis for Automatic Assessment of Pathological Speech Intelligibility
abstract
Speech intelligibility is an important assessment criterion of the communicative performance of pathological speakers. To assist clinicians in their assessment, time- and cost-efficient automatic intelligibility measures offering a repeatable and reliable assessment are desired. In this paper, we propose to automatically assess pathological speech intelligibility based on a distance measure between the subspaces of spectral patterns of the pathological speech signal and of a fully intelligible (healthy) speech signal. To extract the subspace of spectral patterns we investigate two linear decomposition methods, i.e., Principal Component Analysis and Approximate Joint Diagonalization. Pathological speech intelligibility is then derived using a Grassman distance measure which quantifies the difference between the extracted subspaces of pathological and healthy speech. Experiments on an English database of Cerebral Palsy patients show that the proposed intelligibility measure is significantly correlated with subjective intelligibility ratings. In addition, comparisons to state-of-the-art measures show that the proposed subspace-based measure achieves a high performance with a significantly lower computational cost and without imposing any constraints on the speech material of the speakers.
Parvaneh Janbakhshi, Ina Kodrasi, Hervé Bourlard
INTERSPEECH2
2018 Joint Late Reverberation and Noise Power Spectral Density Estimation in a Spatially Homogeneous Noise Field
abstract
Many multi-channel dereverberation and noise reduction techniques such as the multi-channel Wiener filter (MWF) require an estimate of the late reverberation and noise power spectral densities (PSDs). State-of-the-art multi-channel methods for estimating the late reverberation PSD typically assume that the noise PSD matrix is known. Instead of assuming that the noise PSD matrix is known, in this paper we model the noise as a spatially homogeneous sound field with an unknown time-varying PSD and a known time-invariant spatial coherence matrix. Based on this model, two joint estimators of the late reverberation and noise PSDs are proposed, i.e., a non-blocking-based estimator which simultaneously estimates the target signal, late reverberation, and noise PSDs, and a blocking-based estimator which first estimates the late reverberation and noise PSDs at the output of a blocking matrix aiming to block the target signal. Experimental results show that the proposed blocking-based estimator yields the best performance when used in an MWF, even resulting in a similar or better performance than a state-of-the-art blocking-based estimator of the late reverberation PSD which assumes that the noise PSD matrix is known.
Ina Kodrasi, Simon Doclo
ICASSP1
2018 Complexity Reduction of Eigenvalue Decomposition-Based Diffuse Power Spectral Density Estimators Using the Power Method
abstract
In noisy and reverberant environments speech enhancement techniques such as the multi-channel Wiener filter (MWF) can be used to improve speech quality and intelligibility. Assuming that reverberation and ambient noise can be modeled as diffuse sound fields, such techniques require an estimate of the diffuse power spectral density (PSD). Recently a multi-channel diffuse PSD estimator based on the eigenvalue decomposition (EVD) of the prewhitened signal PSD matrix was proposed. The EVD-based PSD estimator is advantageous in comparison to other state-of-the-art PSD estimators, since it does not require knowledge of the relative early transfer functions of the target signal. However, computing the EVD can be computationally expensive, particularly when the number of microphones is large. In this paper we propose to reduce the complexity of the EVD-based PSD estimator by using the iterative power method to compute the eigenvalues. Since the EVD-based PSD estimator only requires the largest eigenvalues, the full EVD is not required and the power method is a well suited computationally efficient technique to estimate these eigenvalues. Experimental results show that using the PSD estimated via the power method in an MWF yields a very similar performance as using the PSD estimated via the full EVD.
Marvin Tammen, Ina Kodrasi, Simon Doclo
ICASSP2
2018 Single-channel Late Reverberation Power Spectral Density Estimation Using Denoising Autoencoders
abstract
In order to suppress the late reverberation in the spectral domain, many single-channel dereverberation techniques rely on an estimate of the late reverberation power spectral density (PSD).In this paper, we propose a novel approach to late reverberation PSD estimation using a denoising autoencoder (DA), which is trained to learn a mapping from the microphone signal PSD to the late reverberation PSD.Simulation results show that the proposed approach yields a high PSD estimation accuracy and generalizes well to unseen data.Furthermore, simulation results show that the proposed DA-based PSD estimate yields a higher PSD estimation accuracy and a similar dereverberation performance than a state-of-the-art statistical PSD estimate, which additionally also requires knowledge of the reverberation time.
Ina Kodrasi, Hervé Bourlard
INTERSPEECH1
2018 Analysis of Eigenvalue Decomposition-Based Late Reverberation Power Spectral Density Estimation
abstract
Many speech dereverberation techniques require an estimate of the late reverberation power spectral density (PSD). State-of-the-art multichannel methods for estimating the late reverberation PSD typically rely on first, an estimate of the relative transfer functions (RTFs) of the target signal; second, a model for the spatial coherence matrix of the late reverberation; and finally, an estimate of the reverberant speech or reverberant and noisy speech PSD matrix. The RTFs, the spatial coherence matrix, and the speech PSD matrix are all prone to modeling and estimation errors in practice, with the RTFs being particularly difficult to estimate accurately, especially in highly reverberant and noisy scenarios. Recently, we proposed an eigenvalue decomposition (EVD)-based late reverberation PSD estimator, which does not require an estimate of the RTFs. In this paper, this EVD-based PSD estimator is further analyzed and its estimation accuracy and computational complexity are analytically compared to a state-of-the-art maximum likelihood (ML) based PSD estimator. It is shown that for perfect knowledge of the RTFs, spatial coherence matrix, and reverberant speech PSD matrix, the ML-based and the EVD-based PSD estimates are both equal to the true late reverberation PSD. In addition, it is shown that for erroneous RTFs but perfect knowledge of the spatial coherence matrix and reverberant speech PSD matrix, the ML-based PSD estimate is larger than or equal to the true late reverberation PSD, whereas the EVD-based PSD estimate is obviously still equal to the true late reverberation PSD. Finally, it is shown that when modeling and estimation errors occur in all quantities, the ML-based PSD estimate is larger than or equal to the EVD-based PSD estimate. Simulation results for several realistic acoustic scenarios demonstrate the advantages of using the EVD-based PSD estimator in a multichannel Wiener filter, yielding a significantly better performance than the ML-based PSD estimator.
Ina Kodrasi, Simon Doclo
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Late reverberant power spectral density estimation based on an eigenvalue decomposition
abstract
Multi-channel methods for estimating the late reverberant power spectral density (PSD) rely on an estimate of the direction of arrival (DOA) of the speech source or of the relative early transfer functions (RETFs) of the target signal from a reference microphone to all microphones. The DOA and the RETFs may be difficult to estimate accurately, particularly in highly reverberant and noisy scenarios. In this paper we propose a novel multi-channel method to estimate the late reverberant PSD which does not require estimates of the DOA or RETFs. The late reverberation is modeled as an isotropic sound field and the late reverberant PSD is estimated based on the eigenvalues of the prewhitened received signal PSD matrix. Experimental results demonstrate the advantages of using the proposed estimator in a multi-channel Wiener filter for speech dereverberation, outperforming a recently proposed maximum likelihood estimator both when the DOA is perfectly estimated as well as in the presence of DOA estimation errors.
Ina Kodrasi, Simon Doclo
ICASSP1
2017 Signal-Dependent Penalty Functions for Robust Acoustic Multi-Channel Equalization
abstract
Acoustic multi-channel equalization techniques, which aim to achieve dereverberation by reshaping the room impulse responses (RIRs) between the source and the microphone array, are known to be highly sensitive to RIR perturbations. In order to increase the robustness against RIR perturbations, several signal-independent methods have been proposed, which only rely on the available perturbed RIRs and do not incorporate any knowledge about the output signal. This paper presents a novel signal-dependent method to increase the robustness of equalization techniques by enforcing the output signal to exhibit spectrotemporal characteristics of a clean speech signal. Motivated by the sparse nature of clean speech, we propose to extend the cost function of state-of-the-art least squares equalization techniques, i.e., the multiple-input/output inverse theorem (MINT), relaxed multi-channel least squares (RMCLS), and partial multi-channel equalization based on MINT (PMINT), with a signal-dependent penalty function promoting sparsity of the output signal in the short-time Fourier transform domain. Three conventionally used sparsity-promoting penalty functions are investigated, i.e., the l0-norm, the l1-norm, and the weighted l1-norm, and the sparsitypromoting reshaping filters are iteratively computed using the alternating direction method of multipliers. Simulation results for several acoustic systems and RIR perturbations demonstrate that incorporating sparsity-promoting penalty functions significantly increases the robustness of MINT, RMCLS, and PMINT, with the weighted l1-norm typically outperforming the l0-norm and the l1-norm. Furthermore, it is shown that the weighted l1-norm sparsity-promoting PMINT technique outperforms the other sparsity-promoting techniques in terms of perceptual speech quality. Finally, it is shown that the signal-dependent weighted l1-norm sparsity-promoting PMINT technique yields a similar or better dereverberation performance than the signal-independent regularized PMINT technique, confirming the advantage of using signal-dependent penalty functions for robust dereverberation filter design.
Ina Kodrasi, Simon Doclo
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Robust sparsity-promoting acoustic multi-channel equalization for speech dereverberation
abstract
This paper presents a novel signal-dependent method to increase the robustness of acoustic multi-channel equalization techniques against room impulse response (RIR) estimation errors. Aiming at obtaining an output signal which better resembles a clean speech signal, we propose to extend the acoustic multi-channel equalization cost function with a penalty function which promotes sparsity of the output signal in the short-time Fourier transform domain. Two conventionally used sparsity-promoting penalty functions are investigated, i.e., the l0-norm and the l1-norm, and the sparsity-promoting filters are iteratively computed using the alternating direction method of multipliers. Simulation results for several RIR estimation errors show that incorporating a sparsity-promoting penalty function significantly increases the robustness, with the l1-norm penalty function outperforming the l0-norm penalty function.
Ina Kodrasi, Ante Jukic, Simon Doclo
ICASSP1
2016 Joint Dereverberation and Noise Reduction Based on Acoustic Multi-Channel Equalization
abstract
Regularized acoustic multi-channel equalization techniques, such as regularized partial multi-channel equalization based on the multiple-input/output inverse theorem (RPMINT), are able to achieve a high dereverberation performance in the presence of room impulse response perturbations but may lead to amplification of the additive noise. In this paper, two time-domain techniques aiming at joint dereverberation and noise reduction based on acoustic multi-channel equalization are proposed. The first technique, namely RPMINT for joint dereverberation and noise reduction (RPM-DNR), extends RPMINT by explicitly taking the noise statistics into account. In addition to the regularization parameter used in RPMINT, the RPM-DNR technique introduces an additional weighting parameter, enabling a trade-off between dereverberation and noise reduction. The second technique, namely multi-channel Wiener filter for joint dereverberation and noise reduction (MWF-DNR), takes both the speech and the noise statistics into account and uses the RPMINT filter to compute a dereverberated reference signal for the multi-channel Wiener filter. The MWF-DNR technique also introduces an additional weighting parameter, which now provides a trade-off between speech distortion and noise reduction. To automatically select the regularization and weighting parameters, for the RPM-DNR technique a novel procedure based on the L-hypersurface is proposed, whereas for the MWF-DNR technique two decoupled optimization procedures based on the L-curve are used. Extensive simulations demonstrate using instrumental measures that the RPM-DNR technique maintains the dereverberation performance of the RPMINT technique while improving its noise reduction performance. Furthermore, it is shown that the MWF-DNR technique yields a significantly better noise reduction performance than the RPM-DNR technique at the expense of a worse dereverberation performance.
Ina Kodrasi, Simon Doclo
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Curvature-based optimization of the trade-off parameter in the speech distortion weighted multichannel wiener filter
abstract
The objective of the speech distortion weighted multichannel Wiener filter (MWF) is to reduce background noise while controlling speech distortion. This can be achieved by means of a trade-off parameter, hence, selecting an optimal trade-off parameter is of crucial importance. Aiming at incorporating knowledge about the resulting speech distortion and noise power, in this paper we propose to compute the trade-off parameter as the point of maximum curvature of the parametric plot of noise power versus speech distortion. To determine a narrowband trade-off parameter, an analytical expression is derived for computing the point of maximum curvature, whereas to determine a broadband parameter an optimization routine is used. The speech distortion and the noise power terms can also be weighted in advance, e.g. based on perceptually motivated criteria. Experimental results show that using the proposed method instead of the MWF improves the intelligibility weighted SNR without significantly degrading the speech distortion.
Ina Kodrasi, Daniel Marquardt, Simon Doclo
ICASSP1
2014 Frequency-domain single-channel inverse filtering for speech dereverberation: Theory and practice
abstract
The objective of single-channel inverse filtering is to design an inverse filter that achieves dereverberation while being robust to an inaccurate room impulse response (RIR) measurement or estimate. Since a stable and causal inverse filter typically does not exist, approximate time-domain inverse filtering techniques such as singlechannel least-squares (SCLS) have been proposed. However, besides being computationally expensive and often infeasible, SCLS generally leads to distortions in the output signal in the presence of RIR inaccuracies. In this paper, a theoretical analysis is initially provided, showing that the direct inversion of the acoustic transfer function in the frequency-domain generally yields instability and acausality issues. In order to resolve these issues, a novel frequency-domain inverse filtering technique is proposed that incorporates regularization and uses a single-channel speech enhancement scheme. Experimental results demonstrate that the proposed technique yields a higher dereverberation performance and has a significantly lower computational complexity compared to the SCLS technique.
Ina Kodrasi, Timo Gerkmann, Simon Doclo
ICASSP1
2013 A perceptually constrained channel shortening technique for speech dereverberation
abstract
The objective of acoustic multichannel equalization is to design a reshaping filter that reduces reverberation, improves the perceptual speech quality, and is robust to errors in the estimated room impulse responses (RIRs). Although the channel shortening (CS) technique has been shown to be effective in achieving dereverberation, it may fail to preserve the natural shape of an RIR leading to speech quality degradation. Furthermore, CS yields multiple reshaping filters that satisfy its optimization criterion but result in a different perceptual speech quality. In this paper, we propose a robust perceptually constrained channel shortening technique (PeCCS) that resolves the selection ambiguity of CS and leads to joint dereverberation and speech quality preservation. Simulation results for erroneously estimated RIRs show that PeCCS preserves the perceptual speech quality and results in a higher reverberant tail suppression than other state-of-the-art techniques, such as CS and the regularized partial multichannel equalization technique based on the multiple-input/output inverse theorem (P-MINT).
Ina Kodrasi, Stefan Goetze, Simon Doclo
ICASSP1
2013 Regularization for Partial Multichannel Equalization for Speech Dereverberation
abstract
Acoustic multichannel equalization techniques such as the multiple-input/output inverse theorem (MINT), which aim to equalize the room impulse responses (RIRs) between the source and the microphone array, are known to be highly sensitive to RIR estimation errors. To increase robustness, it has been proposed to incorporate regularization in order to decrease the energy of the equalization filters. In addition, more robust partial multichannel equalization techniques such as relaxed multichannel least-squares (RMCLS) and channel shortening (CS) have recently been proposed. In this paper, we propose a partial multichannel equalization technique based on MINT (P-MINT) which aims to shorten the RIR. Furthermore, we investigate the effectiveness of incorporating regularization to further increase the robustness of P-MINT and the aforementioned partial multichannel equalization techniques, i.e., RMCLS and CS. In addition, we introduce an automatic non-intrusive procedure for determining the regularization parameter based on the L-curve. Simulation results using measured RIRs show that incorporating regularization in P-MINT yields a significant performance improvement in the presence of RIR estimation errors, whereas a smaller performance improvement is observed when incorporating regularization in RMCLS and CS. Furthermore, it is shown that the intrusively regularized P-MINT technique outperforms all other investigated intrusively regularized multichannel equalization techniques in terms of perceptual speech quality (PESQ). Finally, it is shown that the automatic non-intrusive regularization parameter in regularized P-MINT leads to a very similar performance as the intrusively determined optimal regularization parameter, making regularized P-MINT a robust, perceptually advantageous, and practically applicable multichannel equalization technique for speech dereverberation.
Ina Kodrasi, Stefan Goetze, Simon Doclo
IEEE Trans. Speech Audio Process.1
2012 Robust partial multichannel equalization techniques for speech dereverberation
abstract
This paper presents a novel approach for partial multichannel equalization using the multiple-input/output inverse theorem with the first part of one of the estimated channels as the target response (P-MINT). In order to further increase the robustness against channel estimation errors, two extensions are proposed, i.e. the incorporation of a regularization parameter in the inverse filter design and a truncated singular value decomposition approach. Experimental results for speech dereverberation show that the regularized P-MINT method outperforms state-of-the-art techniques such as channel shortening and the relaxed multichannel least-squares method in terms of robustness to channel estimation errors.
Ina Kodrasi, Simon Doclo
ICASSP1
2011 Microphone position optimization for planar superdirective beamforming
abstract
The performance of a fixed beamformer highly depends on the position of the microphones in the array. In this paper, different heuristic optimisation approaches for arbitrary planar arrays and an exhaustive search approach for structured array geometries are presented to optimise the microphone positions for a superdirective beamformer, aiming at maximizing the mean directivity index for several steering angles of interest. Through the derivation of an upper bound on the achievable performance, it is shown that the proposed approaches generate configurations with a near-optimal performance. In addition, the theoretical results are validated using real measurements, demonstrating the practical usability of the proposed methods.
Ina Kodrasi, Thomas Rohdenburg, Simon Doclo
ICASSP1