Nilesh Madhu

dblp:21/7819 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
12since 2021 · last 2024
0000-0001-9131-3309ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Phase Reconstruction in Single Channel Speech Enhancement Based on Phase Gradients and Estimated Clean-Speech Amplitudes
abstract
Phase gradients can help enforce phase consistency across time and frequency, further improving the output of speech enhancement approaches. Recently, neural networks were used to estimate the phase gradients from the short-term amplitude spectra of clean speech. These were then used to synthesise phase to obtain a plausible time-domain signal. However, using purely synthetic phase in speech enhancement yields unnatural-sounding output. Therefore we derive a closed-form phase estimate that combines the synthetic phase with that of the enhanced speech, yielding more natural output. Secondly, we empirically evaluate the benefit of (re-)training the phase gradient estimation networks on the amplitude spectra of the estimated clean-speech signal. Lastly we apply our proposed phase enhancement to the output of a phase-aware speech enhancement DNN, verifying if an independent phase estimator brings additional advantage. Results show that, compared to the baseline, the proposed approach further improves the DNSMOS scores by ≈ 0.1 on average, and significantly in the first quartile on broadband, quasi-stationary noises, where phase enhancement is expected to have maximum benefit. Training phase gradient estimators on estimated speech spectra is additionally beneficial here. Our method even improves the performance of the phase-aware approach, indicating its feasibility as a generic post-processor for speech enhancement.
Yanjue Song, Nilesh Madhu
ICASSP2
2024 Enhanced Deep Speech Separation in Clustered Ad Hoc Distributed Microphone Environments
abstract
Ad-hoc distributed microphone environments, where microphone locations and numbers are unpredictable, present a challenge to traditional deep learning models, which typically require fixed architectures. To tailor deep learning models to accommodate arbitrary array configurations, the Transform-Average-Concatenate (TAC) layer was previously introduced. In this work, we integrate TAC layers with dual-path transformers for speech separation from two simultaneous talkers in realistic settings. However, the distributed nature makes it hard to fuse information across microphones efficiently. Therefore, we explore the efficacy of blindly clustering microphones around sources of interest prior to enhancement. Experimental results show that this deep cluster-informed approach significantly improves the system's capacity to cope with the inherent variability observed in ad-hoc distributed microphone environments.
Stijn Kindt, Nilesh Madhu, Hong-Goo Kang
INTERSPEECH3
2024 Evaluation of BLE-based audio broadcasting under probabilistic interference
Mathias Baert, Bart Moons, Jowan Pittevils, Yanjue Song, Nilesh Madhu, Jeroen Hoebeke
Comput. Commun.5
2024 Spatially Selective Speaker Separation Using a DNN With a Location Dependent Feature Extraction
abstract
Deep neural networks (DNNs) have proven themselves as an effective means to separate clean speech from noisy mixtures. When there are multiple concurrent talkers, however, unambiguously defining the target output is not trivial, especially if the mixture is single-channel and the talkers are not known in advance. Although this problem can be addressed with permutation invariant training or deep clustering, the performance still suffers in this case. Approaches for compact arrays of multiple microphones can exploit spatial diversity to resolve the ambiguity: a separate output may be generated for each direction of arrival (DOA), or the speaker assignment can be controlled with a location-based training (LBT). Alternatively, we can narrow down the target definition at the input, to perform a spatially selective speaker separation instead of separating all speakers simultaneously. This is achieved by specifying freely adjustable target DOAs. On the one hand, these can be integrated as location-based input features (LBI). On the other hand, the main contribution of this work is a location dependent feature extraction (LDE): we implicitly introduce a DOA dependence in a small part of the DNN by optimizing its parameters for each DOA separately. Experiments demonstrate that LDE outperforms LBT and LBI in terms of instrumental metrics and speech recognition results. A representative audio example is presented for a qualitative impression. An analysis of the spatial selectivity reveals that target and nontarget directions can be distinguished quite well with LDE, which is also verified by recordings of real moving talkers.
Alexander Bohlender, Ann Spriet, Wouter Tirry, Nilesh Madhu
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Improved Deep Speaker Localization and Tracking: Revised Training Paradigm and Controlled Latency
abstract
Even without a separate tracking algorithm, the directions of arrival (DOAs) of moving talkers can be estimated with a deep neural network (DNN) when the movement trajectories used for training allow the generalization to real signals. Previously, we proposed a framework for generating training data with time-variant source activity and sudden DOA changes. Slowly moving sources could be seen as a special case thereof, but were not explicitly modeled. In this paper, we extend this framework by using small jumps between neighboring discrete DOAs to simulate gradual movements. Further, we investigate the benefit of a latency controlled bidirectional recurrent layer in the DNN architecture, such that the required strictly limited context of future frames may still be acceptable for real-time applications. Experiments with real recordings show that the revised data generation leads to more continuous DOA paths, whereas the future context enables a quicker detection of speech onsets and offsets.
Alexander Bohlender, Liesbeth Roelens, Nilesh Madhu
ICASSP3
2023 Exploiting Speaker Embeddings for Improved Microphone Clustering and Speech Separation in ad-hoc Microphone Arrays
abstract
For separating sources captured by ad hoc distributed microphones a key first step is assigning the microphones to the appropriate source-dominated clusters. The features used for such (blind) clustering are based on a fixed length embedding of the audio signals in a high-dimensional latent space. In previous work, the embedding was hand-engineered from the Mel frequency cepstral coefficients and their modulation-spectra. This paper argues that embedding frameworks designed explicitly for the purpose of reliably discriminating between speakers would produce more appropriate features. We propose features generated by the state-of-the-art ECAPA-TDNN speaker verification model for the clustering. We benchmark these features in terms of the subsequent signal enhancement as well as on the quality of the clustering where, further, we introduce 3 intuitive metrics for the latter. Results indicate that in contrast to the hand-engineered features, the ECAPA-TDNN-based features lead to more logical clusters and better performance in the subsequent enhancement stages - thus validating our hypothesis.
Stijn Kindt, Jenthe Thienpondt, Nilesh Madhu
ICASSP3
2023 Aiding Speech Harmonic Recovery in DNN-Based Single Channel Noise Reduction Using Cepstral Excitation Manipulation (CEM) Components
abstract
Weak harmonics of voiced speech segments are often lost during the process of noise suppression – especially at low SNRs. This leads to a distortion in the harmonic structure, and an accompanying loss in quality. In this paper, inspired by previous work on speech harmonic enhancement using statistical methods, we present a loss function component we term cepstral excitation manipulation (CEM) loss, which is constructed based on the fundamental frequency-related cepstral coefficients. This component can be introduced to the training of state-of-the-art architectures and its benefit is benchmarked, here, on CRUSE. Experiments show that the proposed loss function component nicely supplements standard loss functions and the harmonic structure is better preserved. On average, the best system improves by 0.4 on PESQ and 0.47 on DNSMOS compared to the noisy input. Substantial improvements are primarily in low SNRs (-5 dB to 5 dB) – the range for which harmonic recovery is most required.
Yanjue Song, Nilesh Madhu
ICASSP2
2023 Margin-Mixup: A Method for Robust Speaker Verification In Multi-Speaker Audio
abstract
This paper is concerned with the task of speaker verification on audio with multiple overlapping speakers. Most speaker verification systems are designed with the assumption of a single speaker being present in a given audio segment. However, in a real-world setting this assumption does not always hold. In this paper, we demonstrate that current speaker verification systems are not robust against audio with noticeable speaker overlap. To alleviate this issue, we propose margin-mixup, a simple training strategy that can easily be adopted by existing speaker verification pipelines to make the resulting speaker embeddings robust against multi-speaker audio. In contrast to other methods, margin-mixup requires no alterations to regular speaker verification architectures, while attaining better results. On our multi-speaker test set based on VoxCeleb1, the proposed margin-mixup strategy improves the EER on average with 44.4% relative to our state-of-the-art speaker verification baseline systems.
Jenthe Thienpondt, Nilesh Madhu, Kris Demuynck
ICASSP2
2022 Improved Separation of Closely-spaced Speakers by Exploiting Auxiliary Direction of Arrival Information within a U-Net Architecture
abstract
Microphone arrays use spatial diversity for separating concurrent audio sources. Source signals from different directions of arrival (DOAs) are captured with DOA-dependent time-delays between the microphones. These can be exploited in the short-time Fourier transform domain to yield time-frequency masks that extract a target signal while suppressing unwanted components. Using deep neural networks (DNNs) for mask estimation has drastically improved separation performance. However, separation of closely spaced sources remains difficult due to their similar inter-microphone time delays. We propose using auxiliary information on source DOAs within the DNN to improve the separation. This can be encoded by the expected phase differences between the microphones. Alternatively, the DNN can learn a suitable input representation on its own when provided with a multi-hot encoding of the DOAs. Experimental results demonstrate the benefit of this information for separating closely spaced sources.
Stijn Kindt, Alexander Bohlender, Nilesh Madhu
AVSS3
2022 Drone Ego-Noise Cancellation for Improved Speech Capture using Deep Convolutional Autoencoder Assisted Multistage Beamforming
Yanjue Song, Stijn Kindt, Nilesh Madhu
FUSION3
2022 Improved CEM for Speech Harmonic Enhancement in Single Channel Noise Suppression
abstract
The periodic nature of voiced speech is often exploited to restore speech harmonics and to increase inter-harmonic noise suppression. In particular, a recent paper proposed to do this by manipulating the speech harmonic frequencies in the cepstral domain. The manipulations were carried out on the cepstrum of the excitation signal, obtained by the source-filter decomposition of speech. This method was termed Cepstral Excitation Manipulation (CEM). In this contribution we further analyse this method, point out its inherent weakness and propose means to overcome it. First of all, it will be shown by both illustrative examples and theoretical analysis that the existing method underestimates the excitation, especially at low signal to noise ratio (SNR) conditions. This inherent weakness leads to speech harmonic weakening and vocoding due to the insufficient noise suppression in the inter-harmonic regions. Then, we propose two modifications to improve the robustness and performance of CEM in low SNR cases. The first modification is to use an instantaneous amplifying factor adapted to the signal, instead of a pre-defined constant, for the excitation cepstrum. The second modification is to smooth the excitation cepstrum to preserve additional fine structure, instead of discarding it. These modifications result in better preservation of speech harmonics, more refined fine structure and higher inter-harmonic noise suppression. Experimental evaluations using a range of standard instrumental metrics conclusively demonstrate that our proposed modifications clearly outperform the existing method, especially in extremely noisy conditions.
Yanjue Song, Nilesh Madhu
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Exploiting Temporal Context in CNN Based Multisource DOA Estimation
abstract
Supervised learning methods are a powerful tool for direction of arrival (DOA) estimation because they can cope with adverse conditions where simplified models fail. In this work, we consider a previously proposed convolutional neural network (CNN) approach that estimates the DOAs for multiple sources from the phase spectra of the microphones. For speech, specifically, the approach was shown to work well even when trained entirely on synthetically generated data. However, as each frame is processed separately, temporal context cannot be taken into account. This prevents the exploitation of interframe signal correlations, and the fact that DOAs do not change arbitrarily over time. We therefore consider two different extensions of the CNN: the integration of a long short-term memory (LSTM) layer, or of a temporal convolutional network (TCN). In order to accommodate the incorporation of temporal context, the training data generation framework needs to be adjusted. To obtain an easily parameterizable model, we propose to employ Markov chains to realize a gradual evolution of the source activity at different times, frequencies, and directions, throughout a training sequence. A thorough evaluation demonstrates that the proposed configuration for generating training data is suitable for the tasks of single-, and multi-talker localization. In particular, we note that with temporal context, it is important to use speech, or realistic signals in general, for the sources. Experiments with recorded impulse responses and noise reveal that the CNN with the LSTM extension outperforms all other considered approaches, including the plain CNN, and the TCN extension.
Alexander Bohlender, Ann Spriet, Wouter Tirry, Nilesh Madhu
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Least-Squares DOA Estimation with an Informed Phase Unwrapping and Full Bandwidth Robustness
abstract
The weighted least-squares (WLS) direction-of-arrival estimator that minimizes an error based on interchannel phase differences is both computationally simple and flexible. However, the approach has several limitations, including an inability to cope with spatial aliasing and a sensitivity to phase wrapping. The recently proposed phase wrapping robust (PWR)-WLS estimator addresses the latter of these issues, but requires solving a nonconvex optimization problem. In this contribution, we focus on both of the described shortcomings. First, a conceptually simpler alternative to PWR is presented that performs comparably given a good initial estimate. This newly proposed method relies on an unwrapping of the phase differences vector. Secondly, it is demonstrated that all microphone pairs can be utilized at all frequencies with both estimators. When incorporating information from other frequency bins, this permits a localization above the spatial aliasing frequency of the array. Experimental results show that a considerable performance improvement is possible, particularly for arrays with a large microphone spacing.
Alexander Bohlender, Ann Spriet, Wouter Tirry, Nilesh Madhu
ICASSP4
2018 DNN-Supported Speech Enhancement With Cepstral Estimation of Both Excitation and Envelope
abstract
In this paper, we propose and compare various techniques for the estimation of clean spectral envelopes in noisy conditions. The source-filter model of human speech production is employed in combination with a hidden Markov model and/or a deep neural network approach to estimate clean envelope-representing coefficients in the cepstral domain. The cepstral estimators for speech spectral envelope-based noise reduction are both evaluated alone and also in combination with the recently introduced cepstral excitation manipulation (CEM) technique for a priori SNR estimation in a noise reduction framework. Relative to the classical MMSE short time spectral amplitude estimator, we obtain more than 2 dB higher noise attenuation, and relative to our recent CEM technique still 0.5 dB more, in both cases maintaining the quality of the speech component and obtaining considerable SNR improvement.
Samy Elshamy, Nilesh Madhu, Wouter Tirry, Tim Fingscheidt
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Instantaneous A Priori SNR Estimation by Cepstral Excitation Manipulation
abstract
As the a priori signal-to-noise ratio (SNR) contains crucial information about a signal's mixture of speech and noise, its estimation is subject to steady research. In this paper, we introduce a novel a priori SNR estimator based on synthesizing an idealized excitation signal in the cepstral domain. Our approach utilizes a source-filter decomposition in combination with a cepstral excitation manipulation in order to recreate an idealized excitation, which is subsequently shaped by an immanent envelope. In contrast to the well-known decision-directed approach by Ephraim and Malah, an instantaneous estimate is obtained, which is less prone to sudden acoustic environmental changes and musical noise. Additionally, the proposed estimator is able to preserve weak harmonic structures resulting in a spectrum that is more full-bodied. We present both a speaker-independent and a speaker-dependent variant of the new a priori SNR estimator, both showing more than 2 dB ΔSNR improvement versus state of the art, without any significant increase in speech distortion.
Samy Elshamy, Nilesh Madhu, Wouter Tirry, Tim Fingscheidt
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 An iterative speech model-based a priori SNR estimator
Samy Elshamy, Nilesh Madhu, Wouter Tirry, Tim Fingscheidt
INTERSPEECH2
2013 Consistent iterative hard thresholding for signal declipping
abstract
Clipping or saturation in audio signals is a very common problem in signal processing, for which, in the severe case, there is still no satisfactory solution. In such case, there is a tremendous loss of information, and traditional methods fail to appropriately recover the signal. We propose a novel approach for this signal restoration problem based on the framework of Iterative Hard Thresholding. This approach, which enforces the consistency of the reconstructed signal with the clipped observations, shows superior performance in comparison to the state-of-the-art declipping algorithms. This is confirmed on synthetic and on actual high-dimensional audio data processing, both on SNR and on subjective user listening evaluations.
Srdan Kitic, Laurent Jacques, Nilesh Madhu, Michael Peter Hopwood, Ann Spriet, Christophe De Vleeschouwer
ICASSP3
2013 The Potential for Speech Intelligibility Improvement Using the Ideal Binary Mask and the Ideal Wiener Filter in Single Channel Noise Reduction Systems: Application to Auditory Prostheses
abstract
Whereas state-of-the-art single-channel noise reduction algorithms for auditory prostheses demonstrate an appreciable suppression of the noise and improved speech quality, they are unable, thus far, to improve the intelligibility of noise-degraded speech signals. Alternative approaches to speech enhancement using a binary time-frequency mask have demonstrated substantial intelligibility improvements in low signal-to-noise-ratio (SNR) conditions under ideal settings, making this a promising research direction for auditory prostheses. These approaches exploit the sparsity and disjoint-ness of speech spectra in their short-time-frequency representation to preserve only the target-dominant time-frequency regions in the processed output. State-of-the-art noise reduction algorithms in contrast are soft-decision approaches which weight each time-frequency region in proportion to the prevailing SNR. However, the potential for intelligibility improvement using these approaches has not been examined systematically vis-à-vis the binary mask alternative. This contribution compares the performance of an ideal soft-decision system, exemplified by the ideal Wiener filter (IWF), and the ideal binary mask (IBM) for single-channel speech enhancement for auditory prostheses. To obtain results relevant to this application area, a (relatively) low spectral resolution, modelled using the Bark-spectrum scale, is used for both the IWF and the IBM. This spectral resolution is comparable to that being used in commercial hearing instruments. The comparison is in terms of potential for intelligibility improvement and resulting signal quality. Intelligibility tests carried out under various noise conditions and SNRs show that the IWF leads to higher intelligibility scores than the IBM in low SNR conditions. Under non-ideal parameter estimates, it is demonstrated that the IWF approach is also much less sensitive to estimation errors. Quality-wise, a preference for the IWF exists. This was evaluated using a two-stage, pair-wise preference-rating test.
Nilesh Madhu, Ann Spriet, Sofie Jansen, Raphael Koning, Jan Wouters
IEEE Trans. Speech Audio Process.1
2012 Reference Estimation in EEG: Analysis of Equivalent Approaches
abstract
This letter demonstrates the theoretical equivalence between the different ICA/beamforming based solutions proposed in the literature for the reference-estimation problem in electroencephalographic (EEG) recordings. By reference, we understand an unknown, non-null, time-varying potential, measured at the reference electrode situated sufficiently distant from the measuring electrodes. Despite the theoretical equivalence of the various approaches, they do not yield identical results in practice. This discrepancy is primarily due to the practical implementation of the underlying approach. We show in this context that the most reliable solution avoids blind source separation and montage transformation in addition to making full use of available a priori knowledge.
Radu Ranta, Nilesh Madhu
IEEE Signal Process. Lett.2
2011 A Versatile Framework for Speaker Separation Using a Model-Based Speaker Localization Approach
abstract
We build upon our speaker localization framework developed in a previous work (N. Madhu and R. Martin, A scalable framework for multiple speaker localization and tracking,” in Proc. Int. Workshop Acoustic Echo Noise Control (IWAENC), Sep. 2008) to perform source separation. The proposed approach, exploiting the supplementary information from the mixture of Gaussians-based localization model, allows for the incorporation of a wide class of separation algorithms, from the nonlinear time-frequency mask-based approaches to a fully adaptive beamformer in the generalized sidelobe canceller (GSC) structure. We propose, in addition, a generalized estimation of the blocking matrix based on subspace projectors. The adaptive beamformer realized as proposed is insensitive to gain mismatches among the sensors, obviating the need for magnitude calibration of the microphones. It is also demonstrated that the proposed linear approach has a performance comparable to that of an optimal (oracle) GSC implementation. In comparison to ICA-based approaches, another advantage of the separation framework described herein is its robustness to ambient noise and scenarios with an unknown number of sources.
Nilesh Madhu, Rainer Martin 0001
IEEE Trans. Speech Audio Process.1
2008 Temporal smoothing of spectral masks in the cepstral domain for speech separation
abstract
This contribution details the development of a mask-based post- processor to improve the interference suppression in speech signals separated using linear deconvolution algorithms like independent component analysis (ICA). The design of the proposed post-filter is in two stages: in the first stage, use is made of the disjointness of the separated signals in the time-frequency domain to obtain binary masks to suppress cross-talk that generally remains after separation. In the next stage, a novel smoothing of the masks is proposed that preserves the speech structure of the target source while eliminating the random peaks in the time-frequency plane that lead to fluctuating background noise. The result is an enhanced signal with reduced cross-talk and no musical noise.
Nilesh Madhu, Colin Breithaupt, Rainer Martin 0001
ICASSP1
2008 AN EM-based probabilistic approach for Acoustic Echo Suppression
abstract
This paper introduces a new acoustic echo suppression (AES) algorithm for suppressing the residual echo after the acoustic echo canceller (AEC). By temporally segmenting the frequency bins of the residual signal spectrum into blocks and modelling the data in each block and each frequency bin as realizations of a random variable, we can compute the probability of presence of residual echo and derive an appropriate ML suppression rule based on this probability. The computation of the probabilities is based on the Expectation Maximization algorithm. The proposed method shows better performance as compared to state of the art methods for residual echo suppression while producing no audible degradation in the near end signal and no musical noise. Test results indicate that the proposed approach provides an increase in the ERLE of up to 3 dB more than the state of the art echo suppressor while yielding a comparable mean opinion score (MOS) for the near end speech quality. Furthermore, the proposed method is independent of the double talk detector - which makes it robust to misclassifications on the part of the AEC algorithm.
Nilesh Madhu, Ivan Tashev, Alex Acero
ICASSP1
2005 Robust speaker localization through adaptive weighted pair TDOA (AWEPAT) estimation
abstract
Time delay of arrival (TDOA) estimation between signals input to two or more microphones plays an important role in speaker localization.Most methods employ a linear array of two or more microphones and use the generalized cross correlation method or eigenspace analysis (AEDA) methods.TDOA estimation with linear arrays, however, is highly sensitive to estimation errors when the signals arrive from an endfire direction.In this paper we propose a novel adaptive algorithm which makes use of a three-microphone planar array.This algorithm exhibits a much smaller estimation error over the complete azimuth range of 0-360 degrees as compared to other algorithms.The computational complexity of this approach is comparable to other state-of-the-art algorithms.
Nilesh Madhu, Rainer Martin 0001
INTERSPEECH1
2004 Independent component analysis for semi-blind signal separation in MIMO mobile frequency selective communication channels
abstract
In this paper we address the problem of semi-blind source separation (SBSS) in frequency selective MIMO mobile communication channels. Semi-blindness stems from the fact that some average properties of the time-varying channel (mixing domain) are available at the transmitter. In this paper we first analytically show that when orthogonal frequency division multiplexing (OFDM) is employed, the original BSS problem is transformed into a set of standard ICA problems with complex mixing matrices. Each ICA problem is associated with one of the orthogonal subcarriers. This is special case of performing ICA in frequency domain where no inverse Fourier transformation of the separated signals is necessary. Secondly, we show that the statistical correlation between the different frequency bins (at each orthogonal sub-carrier) can be exploited to avoid the frequency dependent permutation problem, intrinsic to the ICA solution. Our approach has been tested on a realistic channel model and the results are presented.
Dragan Obradovic, Nilesh Madhu, Andrei Szabo, Chiu Shun Wong
IJCNN2