Jesper Jensen 0001

dblp:94/356 · DBLP profile ↗
← Back
135ranked-venue papers
17as first author
31since 2021 · last 2026
0000-0003-1478-622XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 85 · 10 first-author · 21 since 2021Artificial intelligence and machine learning · 65 · 10 first-author · 15 since 2021Databases, data management, data science and information retrieval · 2Theory of computation · 2Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TSIP-Net: No-reference speech intelligibility prediction in the presence of competing speech
Haolan Wang, Wai-Yip Chan, Jesper Jensen 0001
Speech Commun.3
2025 Deep Feedback Cancellation for Hearing Aids with Improved System Stability and Sound Quality
abstract
Acoustic feedback cancellation is an important task in audio processing systems, aiming to mitigate the effects of feedback loops on system stability and sound quality. State-of-the-art methods rely on adaptive filtering algorithms and face challenges in balancing between rapid convergence and low steady-state error. In this work, we introduce a novel approach inspired by traditional adaptive filtering and deep learning techniques to achieve a significantly faster convergence and lower steady-state errors at the same time. Our proposed system, termed Deep Feedback Cancellation (DFC), leverages deep neural networks to predict the impulse response of the feedback path directly. Hence, it replaces traditional gradient based adaptive estimation of the impulse responses. Experimental evaluation, in a hearing aid setting, conducted on real-world data demonstrates the superiority of DFC over traditional methods. Specifically, in a practically very important situation, where the feedback path undergoes rapid changes, the proposed DFC achieves increased convergence rate by a factor of 30, while decreasing the steady-state error by 2 dB. Generally, DFC leads to very significant improvements over traditional methods. These improvements are confirmed by objective evaluations and subjective listening tests. Our findings suggest that DFC presents a promising alternative for acoustic feedback cancellation in hearing aid applications.
Eleftheria Lydaki, Zheng-Hua Tan, Jesper Jensen 0001, Meng Guo 0001
ICASSP3
2025 Intelligibility Prediction for Time-Modified Speech Signals Using Spectro-Temporal Modulation Features
abstract
Reference-based speech intelligibility prediction algorithms (RB-SIPAs) are limited to speech degradations that maintain the time alignment between the clean and degraded speech signals. We address this limitation by augmenting the existing RB-SIPAs framework with time alignment. To achieve robust time alignment at low signal-to-noise ratios, we propose using spectro-temporal modulations (STM) of the speech signals for dynamic time warping (DTW). Our experiments demonstrate that in DTW, the use of specific STM components as features for clean and time-modified degraded speech achieves the baseline time alignment obtained using clean speech and noise-free time-modified speech. Moreover, we propose two methods to incorporate time alignment into existing RB-SIPAs. Using these methods, we compare the output scores of the RB-SIPA with listening test scores and show better correlation results using the STM features as compared to MFCCs.
Aymen Bashir, Haolan Wang, Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001
INTERSPEECH5
2025 xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
abstract
While attention-based architectures, such as Conformers, excel in speech enhancement, they face challenges such as scalability with respect to input sequence length. In contrast, the recently proposed Extended Long Short-Term Memory (xLSTM) architecture offers linear scalability. However, xLSTM-based models remain unexplored for speech enhancement. This paper introduces xLSTM-SENet, the first xLSTM-based single-channel speech enhancement system. A direct comparative analysis reveals that xLSTM-and notably, even LSTM-can match or outperform state-of-the-art Mamba- and Conformer-based systems across various model sizes in speech enhancement on the VoiceBank+Demand dataset. Through ablation studies, we identify key architectural design choices such as exponential gating and bidirectionality contributing to its effectiveness. Our best xLSTM-based model, xLSTM-SENet2, outperforms state-of-the-art Mamba- and Conformer-based systems of similar complexity on the Voicebank+DEMAND dataset.
Nikolai Lund Kühne, Jan Østergaard, Jesper Jensen 0001, Zheng-Hua Tan
INTERSPEECH3
2025 Analysis and Extension of a Near-End Listening Enhancement Method Based on Long-Term Fractile Noise Statistics
abstract
This paper addresses the problem of near-end listening enhancement (NELE), where a clean speech signal is modified prior to playback and under an energy constraint to improve intelligibility in noise. We analyze a recently proposed NELE method, optimized using a Speech Intelligibility Index that has been modified to incorporate temporal aspects of the noise via long-term fractile noise statistics. Specifically, we explain the energy allocation strategy adopted by the algorithm, and show that, in contrast to many existing methods, the spectral energy distribution of the modified speech is a function of that of the background noise, but not that of the input speech. Our simulation experiments show that this simple method outperforms well-established spectral shaping NELE methods. In addition, we extend the algorithm by appending an off-the-shelf dynamic range compressor, and show that it performs generally better than state-of-the-art methods for NELE.
Filippo Villani, Wai-Yip Chan, Zheng-Hua Tan, Jan Østergaard, Jesper Jensen 0001
INTERSPEECH5
2025 Noise-Robust Hearing Aid Voice Control
abstract
Advancing the design of robust hearing aid (HA) voice control is crucial to increase the HA use rate among hard of hearing people as well as to improve HA users' experience. In this work, we contribute towards this goal by, first, presenting a novel HA speech dataset consisting of noisy own voice captured by 2 behind-the-ear (BTE) and 1 in-ear-canal (IEC) microphones. Second, we provide baseline HA voice control results from the evaluation of light, state-of-the-art keyword spotting models utilizing different combinations of HA microphone signals. Experimental results show the benefits of exploiting bandwidth-limited bone-conducted speech (BCS) from the IEC microphone to achieve noise-robust HA voice control. Furthermore, results also demonstrate that voice control performance can be boosted by assisting BCS by the broader-bandwidth BTE microphone signals. Aiming at setting a baseline upon which the scientific community can continue to progress, the HA noisy speech dataset has been made publicly available.
Iván López-Espejo, Eros Roselló, Amin Edraki, Naomi Harte, Jesper Jensen 0001
IEEE Signal Process. Lett.5
2024 Self-Supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
abstract
In this paper, we propose the use of self-supervised pretraining on a large unlabelled data set to improve the performance of a personalized voice activity detection (VAD) model in adverse conditions. We pretrain a long short-term memory (LSTM)-encoder using the autoregressive predictive coding (APC) framework and fine-tune it for personalized VAD. We also propose a denoising variant of APC, with the goal of improving the robustness of personalized VAD. The trained models are systematically evaluated on both clean speech and speech contaminated by various types of noise at different SNR-levels and compared to a purely supervised model. Our experiments show that self-supervised pretraining not only improves performance in clean conditions, but also yields models which are more robust to adverse conditions compared to purely supervised learning.
Holger S. Bovbjerg, Jesper Jensen 0001, Jan Østergaard, Zheng-Hua Tan
ICASSP2
2024 Speaker Adaptation For Enhancement Of Bone-Conducted Speech
abstract
Deep neural network (DNN)-based speech enhancement models often face challenges in maintaining their performance for speakers not encountered during training. This challenge is exacerbated in applications such as enhancement and bandwidth extension of bone-conducted speech, where the distortion exhibits a close correlation with speaker-specific characteristics. We address this issue by introducing a bottleneck module aimed at disentangling speaker-specific characteristics from speech content in speech enhancement DNNs. A DNN model is trained for enhancement of bone-conducted speech and modified with the proposed bottleneck module. We evaluate the DNN’s adaptability to unseen speakers through fine-tuning the network with a limited amount of adaptation data. The results show that the proposed bottleneck module can enhance adaptation performance to new unseen speakers, especially when limited amount of speaker-specific adaptation data is available.
Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001, Daniel Fogerty
ICASSP3
2024 Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based Sampler
abstract
Diffusion models are a new class of generative models that have recently been applied to speech enhancement successfully. Previous works have demonstrated their superior performance in mismatched conditions compared to state-of-the art discriminative models. However, this was investigated with a single database for training and another one for testing, which makes the results highly dependent on the particular databases. Moreover, recent developments from the image generation literature remain largely unexplored for speech enhancement. These include several design aspects of diffusion models, such as the noise schedule or the reverse sampler. In this work, we systematically assess the generalization performance of a diffusion-based speech enhancement model by using multiple speech, noise and binaural room impulse response (BRIR) databases to simulate mismatched acoustic conditions. We also experiment with a noise schedule and a sampler that have not been applied to speech enhancement before. We show that the proposed system substantially benefits from using multiple databases for training, and achieves superior performance compared to state-of-the-art discriminative models in both matched and mismatched conditions. We also show that a Heun-based sampler achieves superior performance at a smaller computational cost compared to a sampler commonly used for speech enhancement.
Philippe Gonzalez, Zheng-Hua Tan, Jan Østergaard, Jesper Jensen 0001, Tommy S. Alstrøm, Tobias May
ICASSP4
2024 Speech Enhancement in Hearing Aids Using Target Speech Presence Estimation Based on a Delayed Remote Microphone Signal
abstract
Speech enhancement in hearing aids (HAs) can take advantage of a wireless remote microphone (RM) having a better signal-to-noise ratio than the HA microphones. However, using the RM effectively is complicated by the time delay between the acoustic and wireless signals. Methods in the literature assume an instantaneous transmission of the RM signal, which is never the case in practice. Hence, we propose a practically operational method to use an RM with HAs in the presence of wireless transmission delays. Specifically, we use the delayed target voice activity state from the RM signal, to derive an expression for target speech presence probability (SPP) at the local HA microphone signals. The proposed method uses this target SPP mask as a post-filter that follows a local multichannel Wiener filter. Through simulations, we demonstrate that the proposed method improves speech quality and intelligibility metrics, especially in very noisy acoustic environments, compared to a standard approach which relies solely on HA microphones.
Vasudha Sathyapriyan, Michael Syskind Pedersen, Mike Brookes, Jan Østergaard, Patrick A. Naylor, Jesper Jensen 0001
ICASSP6
2024 Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks
abstract
Studies have shown that in noisy acoustic environments, providing binaural signals to the user of an assistive listening device may improve speech intelligibility and spatial awareness. This paper presents a binaural speech enhancement method using a complex convolutional neural network with an encoder-decoder architecture and a complex multi-head attention transformer. The model is trained to estimate individual complex ratio masks in the time-frequency domain for the left and right-ear channels of binaural hearing devices. The model is trained using a novel loss function that incorporates the preservation of spatial information along with speech intelligibility improvement and noise reduction. Simulation results for acoustic scenarios with a single target speaker and isotropic noise of various types show that the proposed method improves the estimated binaural speech intelligibility and preserves the binaural cues better in comparison with several baseline algorithms.
Vikas Tokala, Eric Grinstein, Mike Brookes, Simon Doclo, Jesper Jensen 0001, Patrick A. Naylor
ICASSP5
2024 No-Reference Speech Intelligibility Prediction Leveraging a Noisy-Speech ASR Pre-Trained Model
abstract
Recent advances in deep learning have improved the capabilities of data-driven speech intelligibility prediction (SIP) algorithms. Nevertheless, the scarcity of speech intelligibility datasets limits the development of data-driven algorithms. This study introduces a set of no-reference SIP algorithms leveraging a pre-trained wav2vec 2.0 backbone. We adapt wav2vec 2.0 for automatic speech recognition under additive noise conditions with a parameter-efficient methodology, low-rank adaptation. We demonstrate no-reference SIP algorithms designed with this approach using a moderate amount of training data. The best designs perform on par or even better than a state-of-the-art reference-based SIP algorithm across a variety of datasets comprising different degradation types.
Haolan Wang, Amin Edraki, Wai-Yip Chan, Iván López-Espejo, Jesper Jensen 0001
INTERSPEECH5
2024 Channel-Configurable Deep Wireless Speech Transmission
abstract
The proliferation of edge-based wireless speech applications necessitates the development of resource-efficient, low-latency speech communication systems capable of functioning across diverse communication channel conditions. Ensuring intelligible speech communication under conditions of constrained resources and low-latency presents a challenging problem within the domain of speech transmission. In this paper, we introduce a very low-latency configurable speech transmission system leveraging joint source-channel coding and deep neural networks (DNNs). Our proposed system is a unified deep neural network system engineered to operate effectively across a wide range of wireless communication channel scenarios. The system encompasses both a joint source-channel encoder and a joint source-channel decoder, each with access to channel state information (CSI). In this context, CSI signifies the type of fading in the wireless channel. Notably, our system has a total latency of 2 ms. Through extensive simulations, we empirically demonstrate that the proposed configurable system closely approximates the performance of ideal systems specifically tailored to individual wireless channel scenarios. Our evaluation is rooted in the assessment of instrumental measures of speech quality and intelligibility, affirming the efficacy of our system in diverse and resource-constrained communication contexts.
Mohammad Bokaei, Jesper Jensen 0001, Simon Doclo, Jan Østergaard
WCNC2
2024 The Effect of Training Dataset Size on Discriminative and Diffusion-Based Speech Enhancement Systems
abstract
The performance of deep neural network-based speech enhancement systems typically increases with the training dataset size. However, studies that investigated the effect of training dataset size on speech enhancement performance did not consider recent approaches, such as diffusion-based generative models. Diffusion models are typically trained with massive datasets for image generation tasks, but whether this is also required for speech enhancement is unknown. Moreover, studies that investigated the effect of training dataset size did not control for the data diversity. It is thus unclear whether the performance improvement was due to the increased dataset size or diversity. Therefore, we systematically investigate the effect of training dataset size on the performance of popular state-of-the-art discriminative and diffusion-based speech enhancement systems in matched conditions. We control for the data diversity by using a fixed set of speech utterances, noise segments and binaural room impulse responses to generate datasets of different sizes. We find that the diffusion-based systems perform the best relative to the discriminative systems in terms of objective metrics with datasets of 10 h or less. However, their objective metrics performance does not improve when increasing the training dataset size as much as the discriminative systems, and they are outperformed by the discriminative systems with datasets of 100 h or more.
Philippe Gonzalez, Zheng-Hua Tan, Jan Østergaard, Jesper Jensen 0001, Tommy S. Alstrøm, Tobias May
IEEE Signal Process. Lett.4
2024 Investigating the Design Space of Diffusion Models for Speech Enhancement
abstract
Diffusion models are a new class of generative models that have shown outstanding performance in image generation literature. As a consequence, studies have attempted to apply diffusion models to other tasks, such as speech enhancement. A popular approach in adapting diffusion models to speech enhancement consists in modelling a progressive transformation between the clean and noisy speech signals. However, one popular diffusion model framework previously laid in image generation literature did not account for such a transformation towards the system input, which prevents from relating the existing diffusion-based speech enhancement systems with the aforementioned diffusion model framework. To address this, we extend this framework to account for the progressive transformation between the clean and noisy speech signals. This allows us to apply recent developments from image generation literature, and to systematically investigate design aspects of diffusion models that remain largely unexplored for speech enhancement, such as the neural network preconditioning, the training loss weighting, the stochastic differential equation (SDE), or the amount of stochasticity injected in the reverse process. We show that the performance of previous diffusion-based speech enhancement systems cannot be attributed to the progressive transformation between the clean and noisy speech signals. Moreover, we show that a proper choice of preconditioning, training loss weighting, SDE and sampler allows to outperform a popular diffusion-based speech enhancement system while using fewer sampling steps, thus reducing the computational cost by a factor of four.
Philippe Gonzalez, Zheng-Hua Tan, Jan Østergaard, Jesper Jensen 0001, Tommy S. Alstrøm, Tobias May
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 How to Train Your Ears: Auditory-Model Emulation for Large-Dynamic-Range Inputs and Mild-to-Severe Hearing Losses
abstract
Advanced auditory models are useful in designing signal-processing algorithms for hearing-loss compensation or speech enhancement. Such auditory models provide rich and detailed descriptions of the auditory pathway, and might allow for individualization of signal-processing strategies, based on physiological measurements. However, these auditory models are often computationally demanding, requiring significant time to compute. To address this issue, previous studies have explored the use of deep neural networks to emulate auditory models and reduce inference time. While these deep neural networks offer impressive efficiency gains in terms of computational time, they may suffer from uneven emulation performance as a function of auditory-model frequency-channels and input sound pressure level, making them unsuitable for many tasks. In this study, we demonstrate that the conventional machine-learning optimization objective used in existing state-of-the-art methods is the primary source of this limitation. Specifically, the optimization objective fails to account for the frequency- and level-dependencies of the auditory model, caused by a large input dynamic range and different types of hearing losses emulated by the auditory model. To overcome this limitation, we propose a new optimization objective that explicitly embeds the frequency- and level-dependencies of the auditory model. Our results show that this new optimization objective significantly improves the emulation performance of deep neural networks across relevant input sound levels and auditory-model frequency channels, without increasing the computational load during inference. Addressing these limitations is essential for advancing the application of auditory models in signal-processing tasks, ensuring their efficacy in diverse scenarios.
Peter Leer, Jesper Jensen 0001, Zheng-Hua Tan, Jan Østergaard, Lars Bramslow
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Data-Driven Non-Intrusive Speech Intelligibility Prediction Using Speech Presence Probability
abstract
Time consuming Speech Intelligibility (SI) listening tests with human subjects can be replaced by algorithmic SI predictors. In recent years, data-driven SI predictors have been showing promising results. A major limiting factor in the advancement of data-driven SI prediction is that there is a scarcity of SI listening test data available to train the data-driven methods. In this article we propose a data-driven SI predictor that does not require access to an underlying noise-free reference signal, i.e.,non-intrusive, and which does not require listening test data for training. Instead, the proposed method exploits a hypothesized link between SI and Speech Presence Probability (SPP). We show that a neural network can be trained on easily obtainable speech in additive noise data to estimate SPP, and that a simple post-processing stage can be applied in order to map the estimated SPP to SI predictions with high accuracy. The proposed method is evaluated and compared to other state-of-the art non-intrusive SI predictors, and achieves the highest performance even in the presence of processed noisy speech, which the SPP estimator has not been trained on.
Mathias Bach Pedersen, Søren Holdt Jensen, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Distributed Adaptive Norm Estimation for Blind System Identification in Wireless Sensor Networks
abstract
Distributed signal-processing algorithms in (wireless) sensor networks often aim to decentralize processing tasks to reduce communication cost and computational complexity or avoid reliance on a single device (i.e., fusion center) for processing. In this contribution, we extend a distributed adaptive algorithm for blind system identification that relies on the estimation of a stacked network-wide consensus vector at each node, the computation of which requires either broadcasting or relaying of node-specific values (i.e., local vector norms) to all other nodes. The extended algorithm employs a distributed-averaging-based scheme to estimate the network-wide consensus norm value by only using the local vector norm provided by neighboring sensor nodes. We introduce an adaptive mixing factor between instantaneous and recursive estimates of these norms for adaptivity in a time-varying system. Simulation results show that the extension provides estimation results close to the optimal fully-connected-network or broadcasting case while reducing inter-node transmission significantly.
Matthias Blochberger, Filip Elvander, Randall Ali, Jan Østergaard, Jesper Jensen 0001, Marc Moonen, Toon van Waterschoot
ICASSP5
2023 Filterbank Learning for Noise-Robust Small-Footprint Keyword Spotting
abstract
In the context of keyword spotting (KWS), the replacement of handcrafted speech features by learnable features has not yielded superior KWS performance. In this study, we demonstrate that filterbank learning outperforms handcrafted speech features for KWS whenever the number of filterbank channels is severely decreased. Reducing the number of channels might yield certain KWS performance drop, but also a substantial energy consumption reduction, which is key when deploying common always-on KWS on low-resource devices. Experimental results on a noisy version of the Google Speech Commands Dataset show that filterbank learning adapts to noise characteristics to provide a higher degree of robustness to noise, especially when dropout is integrated. Thus, switching from typically used 40-channel log-Mel features to 8channel learned features leads to a relative KWS accuracy loss of only 3.5% while simultaneously achieving a 6.3× energy consumption reduction.
Iván López-Espejo, Ram C. M. C. Shekar, Zheng-Hua Tan, Jesper Jensen 0001, John H. L. Hansen
ICASSP4
2023 Speech inpainting: Context-based speech synthesis guided by video
abstract
Audio and visual modalities are inherently connected in speech signals: lip movements and facial expressions are correlated with speech sounds. This motivates studies that incorporate the visual modality to enhance an acoustic speech signal or even restore missing audio information. Specifically, this paper focuses on the problem of audio-visual speech inpainting, which is the task of synthesizing the speech in a corrupted audio segment in a way that it is consistent with the corresponding visual content and the uncorrupted audio context. We present an audio-visual transformer-based deep learning model that leverages visual cues that provide information about the content of the corrupted audio. It outperforms the previous state-of-the-art audio-visual model and audio-only baselines. We also show how visual features extracted with AV-HuBERT, a large audiovisual transformer for speech recognition, are suitable for synthesizing speech.
Juan F. Montesinos, Daniel Michelsanti, Gloria Haro, Zheng-Hua Tan, Jesper Jensen 0001
INTERSPEECH5
2023 On the deficiency of intelligibility metrics as proxies for subjective intelligibility
abstract
A recent trend in deep neural network (DNN)-based speech enhancement consists of using intelligibility and quality metrics as loss functions for model training with the aim of achieving high subjective speech intelligibility and perceptual quality in real-life conditions. In this study, we analyze a variety of loss functions, including some based on state-of-the-art intelligibility and quality metrics, to train an end-to-end speech enhancement system based on a fully convolutional neural network. The loss functions include perceptual metric for speech quality evaluation (PMSQE), scale-invariant signal-to-distortion ratio (SI-SDR), SI-SDR integrating speech pre-emphasis, short-time objective intelligibility (STOI), extended STOI (ESTOI), spectro-temporal glimpsing index (STGI), and a composite loss function combining STGI and SI-SDR. While DNNs trained with these loss functions produce notable speech intelligibility (and quality) gains according to pertinent objective metrics, we conduct a subjective intelligibility test that contradicts this result, showing no intelligibility improvement. From the results of this study, our conclusion is twofold: (1) subjective intelligibility evaluation is currently not replaceable by objective intelligibility evaluation, and (2) both the development of meaningful intelligibility metrics and DNN-based speech enhancement systems that can consistently improve the intelligibility of noisy speech for human listening remain open problems.
Iván López-Espejo, Amin Edraki, Wai-Yip Chan, Zheng-Hua Tan, Jesper Jensen 0001
Speech Commun.5
2023 Minimum Processing Near-End Listening Enhancement
abstract
The intelligibility and quality of speech from a mobile phone or public announcement system are often affected by background noise in the listening environment. By pre-processing the speech signal it is possible to improve the speech intelligibility and quality — this is known as near-end listening enhancement (NLE). Although, existing NLE techniques are able to greatly increase intelligibility in harsh noise environments, in favorable noise conditions the intelligibility of speech reaches a ceiling where it cannot be further enhanced. Actually, the focus of existing methods solely on improving the intelligibility causes unnecessary processing of the speech signal and leads to speech distortions and quality degradations. In this article, we provide a new rationale for NLE, where the target speech is minimally processed in terms of a processing penalty, provided that a certain performance constraint, e.g., intelligibility, is satisfied. We present a closed-form solution for the case where the performance criterion is an intelligibility estimator based on the approximated speech intelligibility index and the processing penalty is the mean-square error between the processed and the clean speech. This produces an NLE method that adapts to changing noise conditions via a simple gain rule by limiting the processing to the minimum necessary to achieve a desired intelligibility, while at the same time focusing on quality in favorable noise situations by minimizing the amount of speech distortions. Through simulation studies, we show the proposed method attains speech quality on par or better than existing methods in both objective measurements and subjective listening tests, whilst still sustaining objective speech intelligibility performance on par with existing methods.
Andreas Jonas Fuglsig, Jesper Jensen 0001, Zheng-Hua Tan, Lars Søndergaard Bertelsen, Jens Christian Lindof, Jan Østergaard
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Joint Far- and Near-End Speech Intelligibility Enhancement Based on the Approximated Speech Intelligibility Index
abstract
This paper considers speech enhancement of signals picked up in one noisy environment which must be presented to a listener in another noisy environment. Recently, it has been shown that an optimal solution to this problem requires the consideration of the noise sources in both environments jointly. However, the existing optimal mutual information based method requires a complicated system model that includes natural speech variations, and relies on approximations and assumptions of the underlying signal distributions. In this paper, we propose to use a simpler signal model and optimize speech intelligibility based on the Approximated Speech Intelligibility Index (ASII). We derive a closed-form solution to the joint far- and near-end speech enhancement problem that is independent of the marginal distribution of signal coefficients, and that achieves similar performance to existing work. In addition, we do not need to model or optimize for natural speech variations.
Andreas Jonas Fuglsig, Jan Østergaard, Jesper Jensen 0001, Lars Søndergaard Bertelsen, Peter Mariager, Zheng-Hua Tan
ICASSP3
2022 Multichannel Speech Enhancement With Own Voice-Based Interfering Speech Suppression for Hearing Assistive Devices
abstract
Enhancementof a desired speech signal in the presence of competing or interfering speech remains an unsolved problem, as it can be hard to determine which of the speech signals is the one of interest. In this paper, we propose a multichannel noise reduction algorithm which uses the presence of the user’s own voice signal, e.g. during conversations with the target speaker, as an asset to efficiently identify interfering speech and noise. Specifically, following the typical speech pattern in natural conversations, the presence of an own voice may indicate the absence of the target speech, hence undesired speech and noise can be identified and estimated during own voice presence. In contrast to conventional noise reduction systems, the proposed noise reduction systems use the user’s own voice to identify interfering speech that otherwise could be confused with the target speech. We demonstrate the performance of the proposed noise reduction systems in a comparison against state-of-the-art noise reduction systems in terms of beamforming performance for hearing assistive devices. The results show that the proposed beamforming scheme in particular outperforms state-of-the-art methods in terms of ESTOI and PESQ in situations with a target speaker and a strong interfering speaker.
Poul Hoang, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Joint Maximum Likelihood Estimation of Power Spectral Densities and Relative Acoustic Transfer Functions for Acoustic Beamforming
abstract
Acoustic beamforming is crucial for many applications where ex-traction of a target signal from a noisy environment is required. In order to implement practical beamformers, e.g. the multichannel Wiener filter (MWF), estimation of the target and noise power spectral densities (PSDs), and the relative acoustic transfer functions (RATFs) is essential. Several methods, e.g. the so-called covariance whitening (CW) approach, have been proposed for estimating these parameters. However, it seems largely unknown that the CW approach in fact leads to maximum likelihood (ML) estimates of the RATFs. We use historical results to derive joint ML estimates (MLEs) of the RATFs and PSDs in the context of acoustic beam-forming. In addition, based on the MLEs, we propose a basic VAD framework using concentrated likelihood ratios. We use the joint MLEs of the PSDs, RATFs, and the proposed VAD to implement beamformers in a hearing aid application, and compare its performance to competing methods. Simulation results show that the pro-posed scheme can outperform competing methods, in particular in realistic situations where highly accurate prior RATF knowledge is not available or at higher signal-to-noise ratios.
Poul Hoang, Zheng-Hua Tan, Jan Mark de Haan, Jesper Jensen 0001
ICASSP4
2021 Audio-Visual Speech Inpainting with Deep Learning
abstract
In this paper, we present a deep-learning-based framework for audio-visual speech inpainting, i.e., the task of restoring the missing parts of an acoustic speech signal from reliable audio context and uncorrupted visual information. Recent work focuses solely on audio-only methods and generally aims at inpainting music signals, which show highly different structure than speech. Instead, we inpaint speech signals with gaps ranging from 100 ms to 1600 ms to investigate the contribution that vision can provide for gaps of different duration. We also experiment with a multi-task learning approach where a phone recognition task is learned together with speech inpainting. Results show that the performance of audio-only speech inpainting approaches degrades rapidly when gaps get large, while the proposed audio-visual approach is able to plausibly restore missing information. In addition, we show that multi-task learning is effective, although the largest contribution to performance comes from vision.
Giovanni Morrone, Daniel Michelsanti, Zheng-Hua Tan, Jesper Jensen 0001
ICASSP4
2021 A Spectro-Temporal Glimpsing Index (STGI) for Speech Intelligibility Prediction
abstract
We propose a monaural intrusive speech intelligibility prediction (SIP) algorithm called STGI based on detecting glimpses in short-time segments in a spectro-temporal modulation decomposition of the input speech signals. Unlike existing glimpse-based SIP methods, the application of STGI is not limited to additive uncorrelated noise; STGI can be employed in a broad range of degradation conditions. Our results show that STGI performs consistently well across 15 datasets covering degradation conditions including modulated noise, noise reduction processing, reverberation, near-end listening enhancement, checkerboard noise, and gated noise.
Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001, Daniel Fogerty
Interspeech3
2021 Speech Intelligibility Prediction Using Spectro-Temporal Modulation Analysis
abstract
Spectro-temporal modulations are believed to mediate the analysis of speech sounds in the human primary auditory cortex. Inspired by humans' robustness in comprehending speech in challenging acoustic environments, we propose an intrusive speech intelligibility prediction (SIP) algorithm, wSTMI, for normal-hearing listeners based on spectro-temporal modulation analysis (STMA) of the clean and degraded speech signals. In the STMA, each of 55 modulation frequency channels contributes an intermediate intelligibility measure. A sparse linear model with parameters optimized using Lasso regression results in combining the intermediate measures of 8 of the most salient channels for SIP. In comparison with a suite of 10 SIP algorithms, wSTMI performs consistently well across 13 datasets, which together cover degradation conditions including modulated noise, noise reduction processing, reverberation, near-end listening enhancement, and speech interruption. We show that the optimized parameters of wSTMI may be interpreted in terms of modulation transfer functions of the human auditory system. Thus, the proposed approach offers evidence affirming previous studies of the perceptual characteristics underlying speech signal intelligibility.
Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001, Daniel Fogerty
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 A Novel Loss Function and Training Strategy for Noise-Robust Keyword Spotting
abstract
The development of keyword spotting (KWS) systems that are accurate in noisy conditions remains a challenge. Towards this goal, in this paper we propose a novel training strategy relying on multi-condition training for noise-robust KWS. By this strategy, we think of the state-of-the-art KWS models as the composition of a keyword embedding extractor and a linear classifier that are successively trained. To train the keyword embedding extractor, we also propose a new (CN,2+1)-pair loss function extending the concept behind related loss functions like triplet and N-pair losses to reach larger inter-class and smaller intra-class variation. Experimental results on a noisy version of the Google Speech Commands Dataset show that our proposal achieves around 12% KWS accuracy relative improvement with respect to standard end-to-end multi-condition training when speech is distorted by unseen noises. This performance improvement is achieved without increasing the computational complexity of the KWS model.
Iván López-Espejo, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation
abstract
Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, these tasks have been tackled using signal processing and machine learning techniques applied to the available acoustic signals. Since the visual aspect of speech is essentially unaffected by the acoustic environment, visual information from the target speakers, such as lip movements and facial expressions, has also been used for speech enhancement and speech separation systems. In order to efficiently fuse acoustic and visual information, researchers have exploited the flexibility of data-driven approaches, specifically deep learning, achieving strong performance. The ceaseless proposal of a large number of techniques to extract features and fuse multimodal information has highlighted the need for an overview that comprehensively describes and discusses audio-visual speech enhancement and separation based on deep learning. In this paper, we provide a systematic survey of this research topic, focusing on the main elements that characterise the systems in the literature: acoustic features; visual features; deep learning methods; fusion techniques; training targets and objective functions. In addition, we review deep-learning-based methods for speech reconstruction from silent videos and audio-visual sound source separation for non-speech signals, since these methods can be more or less directly applied to audio-visual speech enhancement and separation. Finally, we survey commonly employed audio-visual speech datasets, given their central role in the development of data-driven approaches, and evaluation methods, because they are generally used to compare different systems and determine their performance.
Daniel Michelsanti, Zheng-Hua Tan, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.7
2021 Minimum Processing Beamforming
abstract
Most of the well-known classic beamformers have resulted from optimization problems that minimize a cost function such as the mean-square error (MSE) between the noisy speech and a reference clean speech. The rationale behind these formulations involves a speech-versus-noise dichotomy, where anything branded as noise shall be suppressed as much as possible. While leading to simple closed-form solutions and reasonably practical beamformers, this rationale has its own limitations, for instance, when the ambient noise provides context and is therefore not entirely undesirable. In this article, we offer a new rationale, where the output of the beamformer is minimally processed with respect to a certain reference signal, as long as a given performance criterion is fulfilled. We provide a case study where the performance criterion is inspired by the Speech Intelligibility Index (SII), and the processing penalty is MSE. Regarding the reference signal, we consider two cases. In the first case, the reference signal is set to the unprocessed recording from a reference microphone, giving rise to a beamformer that limits the processing of the noisy signal to a minimum necessary for fulfilling the intelligibility requirement. For the second case, the reference signal is the output of an aggressive beamformer, yielding a beamformer that essentially eliminates the noise unless the concomitant distortion of the clean speech violates the intelligibility requirement. Through simulation studies, we demonstrate some of the benefits that each of the two cases offer in relevant contexts.
Adel Zahedi, Michael Syskind Pedersen, Jan Østergaard, Thomas Ulrich Christiansen, Lars Bramslow, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2020 Maximum Likelihood Estimation of the Interference-Plus-Noise Cross Power Spectral Density Matrix for Own Voice Retrieval
abstract
In headset and hearing aid applications, it is of interest to retrieve the user's own voice in a noisy environment, e.g. for telephony applications. To do so, the cross-power spectral density (CPSD) of the interference-plus-noise is required. In this paper, a novel maximum likelihood (ML) estimator of the interference-plus-noise CPSD matrix is proposed. The proposed method is able to estimate the interference-plus-noise CPSD matrix, even during signal regions with own voice activity. The method uses a novel procedure for estimating the interference-plus-noise CPSD matrix by first estimating the interference PSD and afterwards the noise PSD in a maximum likelihood sense. Simulation experiments, where the proposed method is compared to other noise CPSD matrix estimators, show that it performs on par or better than competing methods, particularly, in situation where the interferenceto-noise ratio is large.
Poul Hoang, Zheng-Hua Tan, Thomas Lunner, Jan Mark de Haan, Jesper Jensen 0001
ICASSP5
2020 A Neural Network for Monaural Intrusive Speech Intelligibility Prediction
abstract
Monaural intrusive speech intelligibility prediction (SIP) methods aim to predict the speech intelligibility (SI) of a single-microphone noisy and/or processed speech signal using the underlying clean speech signal. In the present work, we propose a neural network for monaural intrusive SIP. The proposed network is trained on data from multiple listening tests to predict SI. In the interest of using the available listening test data as efficiently as possible and to facilitate SI prediction of short duration speech signals, training is based on a local-time intelligibility curve derived from the listening test data. The trained neural network is evaluated, in terms of rank order correlation, against the classical monaural intrusive predictors STOI and ESTOI. The network is found to perform the best overall with a Kendall's tau of 0.825 measured over long duration, i.e. speech signals up to several minutes in duration. For short-term prediction using short speech signals of 1 - 10 seconds the network also shows better performance and smaller prediction variance.
Mathias Bach Pedersen, Asger Heidemann Andersen, Søren Holdt Jensen, Jesper Jensen 0001
ICASSP4
2020 A Constrained Maximum Likelihood Estimator of Speech and Noise Spectra with Application to Multi-Microphone Noise Reduction
abstract
One of the challenges with the implementation of multi-microphone noise reduction systems in practical applications lies in the need for the knowledge of the speech and noise covariance matrices. Recently, a method based on Maximum Likelihood (ML) estimation addressed this problem. Despite its relative success in practical setups, this method may suggest negative spectral components for the clean speech due to noise influences. In this paper, we suggest a new estimation technique that tackles this issue by enforcing a power constraint on the estimation problem. We compare the proposed method with the ML method both in synthetic and real-life scenarios using objective measures. The results suggest that the proposed method can improve speech quality without a loss of intelligibility.
Adel Zahedi, Michael Syskind Pedersen, Jan Østergaard, Lars Bramslow, Thomas Ulrich Christiansen, Jesper Jensen 0001
ICASSP6
2020 Vocoder-Based Speech Synthesis from Silent Videos
abstract
Both acoustic and visual information influence human perception of speech. For this reason, the lack of audio in a video sequence determines an extremely low speech intelligibility for untrained lip readers. In this paper, we present a way to synthesise speech from the silent video of a talker using deep learning. The system learns a mapping function from raw video frames to acoustic features and reconstructs the speech with a vocoder synthesis algorithm. To improve speech reconstruction performance, our model is also trained to predict text information in a multi-task learning fashion and it is able to simultaneously reconstruct and recognise speech in real time. The results in terms of estimated speech quality and intelligibility show the effectiveness of our method, which exhibits an improvement over existing video-to-speech approaches.
Daniel Michelsanti, Olga Slizovskaia, Gloria Haro, Emilia Gómez, Zheng-Hua Tan, Jesper Jensen 0001
INTERSPEECH6
2020 End-to-End Speech Intelligibility Prediction Using Time-Domain Fully Convolutional Neural Networks
abstract
Data-driven speech intelligibility prediction has been slow totake off. Datasets of measured speech intelligibility are scarce,and so current models are relatively small and rely on hand-picked features. Classical predictors based on psychoacousticmodels and heuristics are still the state-of-the-art. This workproposes a U-Net inspired fully convolutional neural networkarchitecture, NSIP, trained and tested on ten datasets to pre-dict intelligibility of time-domain speech. The architecture iscompared to a frequency domain data-driven predictor and tothe classical state-of-the-art predictors STOI, ESTOI, HASPIand SIIB. The performance of NSIP is found to be superior fordatasets seen in the training phase. On unseen datasets NSIPreaches performance comparable to classical predictors.
Mathias Bach Pedersen, Morten Kolbæk, Asger Heidemann Andersen, Søren Holdt Jensen, Jesper Jensen 0001
INTERSPEECH5
2020 Spatially Correct Rate-Constrained Noise Reduction for Binaural Hearing Aids in Wireless Acoustic Sensor Networks
abstract
Compared to monaural hearing aids (HAs), binaural hearing aid systems, in which there is a communication link between the two devices, have improved noise reduction capabilities and the ability to preserve binaural spatial information. However, the limited HA battery lifetime puts constraints on the amount of information that can be shared between the two devices. In other words, the rate of transmission between the devices is an important constraint that needs to be considered, while preserving the spatial information. In this article, a linearly constrained noise reduction problem is proposed, which jointly finds the optimal rate allocation and the optimal estimation (beamforming) weights across all sensors and frequencies, while preserving the binaural spatial cues of point sources. The proposed method considers a rate constraint together with linear constraints to preserve the binaural spatial cues of point sources. Minimizing the mean square error on the estimated target speech at the left and the right side beamformers, the optimal weights are found to be rate-constrained linearly constrained minimum variance (LCMV) filters, and the optimal rates are found to be the solutions to a set of reverse water filling problems. The performance of the proposed method is evaluated using the averaged binaural signal-to-noise ratio (SNR), the interaural level difference (ILD) error and the interaural time difference (ITD) error. The results show that the proposed method outperforms spatially correct noise reduction approaches that use naive/random rate allocation strategies.
Jamal Amini, Richard C. Hendriks, Richard Heusdens, Meng Guo 0001, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2020 Rate-Constrained Noise Reduction in Wireless Acoustic Sensor Networks
abstract
Wireless acoustic sensor networks (WASNs) can be used for centralized multi-microphone noise reduction, where the processing is done in a fusion center (FC). To perform the noise reduction, the data needs to be transmitted to the FC. Considering the limited battery life of the devices in a WASN, the total data rate at which the FC can communicate with the different network devices should be constrained. In this article, we propose a rate-constrained multi-microphone noise reduction algorithm, which jointly finds the best rate allocation and estimation weights for the microphones across all frequencies. The optimal linear estimators are found to be the quantized Wiener filters, and the rates are the solutions to a filter-dependent reverse water-filling problem. The performance of the proposed framework is evaluated using simulations in terms of mean square error and predicted speech intelligibility. The results show that the proposed method is very close in performance to that of the existing optimal method based on discrete optimization. However, the proposed approach can do this at a much lower complexity, while the existing optimal reference method needs a non-tractable exhaustive search to find the best rate allocation across microphones.
Jamal Amini, Richard C. Hendriks, Richard Heusdens, Meng Guo 0001, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2020 On Loss Functions for Supervised Monaural Time-Domain Speech Enhancement
abstract
Many deep learning-based speech enhancement algorithms are designed to minimize the mean-square error (MSE) in some transform domain between a predicted and a target speech signal. However, optimizing for MSE does not necessarily guarantee high speech quality or intelligibility, which is the ultimate goal of many speech enhancement algorithms. Additionally, only little is known about the impact of the loss function on the emerging class of time-domain deep learning-based speech enhancement systems. We study how popular loss functions influence the performance of time-domain deep learning-based speech enhancement systems. First, we demonstrate that perceptually inspired loss functions might be advantageous over classical loss functions like MSE. Furthermore, we show that the learning rate is a crucial design parameter even for adaptive gradient-based optimizers, which has been generally overlooked in the literature. Also, we found that waveform matching performance metrics must be used with caution as they in certain situations can fail completely. Finally, we show that a loss function based on scale-invariant signal-to-distortion ratio (SI-SDR) achieves good general performance across a range of popular speech enhancement evaluation metrics, which suggests that SI-SDR is a good candidate as a general-purpose loss function for speech enhancement systems.
Morten Kolbæk, Zheng-Hua Tan, Søren Holdt Jensen, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Improved External Speaker-Robust Keyword Spotting for Hearing Assistive Devices
abstract
For certain applications, keyword spotting (KWS) requires some degree of personalization. This is the case for KWS for hearing assistive devices, e.g., hearing aids, where only the device user should be allowed to trigger the KWS system. In this paper, we first develop a new realistic hearing aid experimental framework. Next, using this framework we show that the performance of a state-of-the-art multi-task deep learning architecture exploiting cepstral features for joint KWS and users' own-voice/external speaker detection drops significantly. To overcome this problem, we use phase difference information through GCC-PHAT (Generalized Cross-Correlation with PHAse Transform)-based coefficients along with log-spectral magnitude features. In addition, we demonstrate that working in the perceptually-motivated constant-Q transform (CQT) domain instead of in the short-time Fourier transform (STFT) domain allows for the generation of compact and coherent features which provide superior KWS performance. Our experimental results show that our CQT-based proposal achieves a relative KWS accuracy improvement of around 18% compared to using cepstral features while dramatically decreasing the number of multiplications in the multi-task architecture, which is key in the context of low-resource devices like hearing assistive devices.
Iván López-Espejo, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Online Multichannel Speech Enhancement Based on Recursive EM and DNN-Based Speech Presence Estimation
abstract
This article presents a recursive expectation-maximization algorithm for online multichannel speech enhancement. A deep neural network mask estimator is used to compute the speech presence probability, which is then improved by means of statistical spatial models of the noisy speech and noise signals. The clean speech signal is estimated using beamforming, single-channel linear postfiltering and speech presence masking. The clean speech statistics and speech presence probabilities are finally used to compute the acoustic parameters for beamforming and postfiltering by means of maximum likelihood estimation. This iterative procedure is carried out on a frame-by-frame basis. The algorithm integrates the different estimates in a common statistical framework suitable for online scenarios. Moreover, our method can successfully exploit spectral, spatial and temporal speech properties. Our proposed algorithm is tested in different noisy environments using the multichannel recordings of the CHiME-4 database. The experimental results show that our method outperforms other related state-of-the-art approaches in noise reduction performance, while allowing low-latency processing for real-time applications.
Juan M. Martín-Doñas, Jesper Jensen 0001, Zheng-Hua Tan, Ángel M. Gómez, Antonio M. Peinado
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 A Novel Binaural Beamforming Scheme with Low Complexity Minimizing Binaural-cue Distortions
abstract
While the majority of binaural beamformers aim to minimize the output noise power while (approximately) preserving the binaural cues of the sources using constraints, we propose in this paper to minimize the binaural-cue distortions of the sources in the acoustic scene, such that the output noise power is below a predefined threshold. This new problem formulation is a convex QCQP problem, which leads to an efficient trade-off between noise reduction, binaural-cue preservation and complexity. In particular, the proposed beamformer provides a better trade-off between noise reduction and binaural-cue preservation (in terms of interaural level and phase differences) compared to the well-known binaural minimum variance distortionless response-η beamformer.
Andreas I. Koutrouvelis, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001, Meng Guo 0001
ICASSP4
2019 Effects of Lombard Reflex on the Performance of Deep-learning-based Audio-visual Speech Enhancement Systems
abstract
Humans tend to change their way of speaking when they are immersed in a noisy environment, a reflex known as Lombard effect. Current speech enhancement systems based on deep learning do not usually take into account this change in the speaking style, because they are trained with neutral (non-Lombard) speech utterances recorded under quiet conditions to which noise is artificially added. In this paper, we investigate the effects that the Lombard reflex has on the performance of audio-visual speech enhancement systems based on deep learning. The results show that a gap in the performance of as much as approximately 5 dB between the systems trained on neutral speech and the ones trained on Lombard speech exists. This indicates the benefit of taking into account the mismatch between neutral and Lombard speech in the design of audio-visual speech enhancement systems.
Daniel Michelsanti, Zheng-Hua Tan, Sigurður Sigurðsson, Jesper Jensen 0001
ICASSP4
2019 On Training Targets and Objective Functions for Deep-learning-based Audio-visual Speech Enhancement
abstract
Audio-visual speech enhancement (AV-SE) is the task of improving speech quality and intelligibility in a noisy environment using audio and visual information from a talker. Recently, deep learning techniques have been adopted to solve the AV-SE task in a supervised manner. In this context, the choice of the target, i.e. the quantity to be estimated, and the objective function, which quantifies the quality of this estimate, to be used for training is critical for the performance. This work is the first that presents an experimental study of a range of different targets and objective functions used to train a deep-learning-based AV-SE system. The results show that the approaches that directly estimate a mask perform the best overall in terms of estimated speech quality and intelligibility, although the model that directly estimates the log magnitude spectrum performs as good in terms of estimated speech quality.
Daniel Michelsanti, Zheng-Hua Tan, Sigurður Sigurðsson, Jesper Jensen 0001
ICASSP4
2019 Improvement and Assessment of Spectro-Temporal Modulation Analysis for Speech Intelligibility Estimation
abstract
Several recent high-performing intelligibility estimators of acoustically degraded speech signals employ temporal modulation analysis. In this paper, we investigate the utility of using both spectro- and temporal-modulation for estimating speech intelligibility. We modified a pre-existing speech intelligibility estimation scheme (STMI) that was inspired by human auditory spectro-temporal modulation analysis. We produced several variants of the modified STMI and assessed their intelligibility prediction accuracy, in comparison with several high-performing estimators. Among the estimators tested, one of the STMI variants and eSTOI performed consistently well on both noisy and reverberated speech. These results suggest that spectro-temporal modulation analysis is useful for certain degradation conditions such as modulated noise and reverberation.
Amin Edraki, Wai-Yip Chan, Jesper Jensen 0001, Daniel Fogerty
INTERSPEECH3
2019 Keyword Spotting for Hearing Assistive Devices Robust to External Speakers
abstract
Keyword spotting (KWS) is experiencing an upswing due to the pervasiveness of small electronic devices that allow interaction with them via speech. Often, KWS systems are speaker-independent, which means that any person --user or not-- might trigger them. For applications like KWS for hearing assistive devices this is unacceptable, as only the user must be allowed to handle them. In this paper we propose KWS for hearing assistive devices that is robust to external speakers. A state-of-the-art deep residual network for small-footprint KWS is regarded as a basis to build upon. By following a multi-task learning scheme, this system is extended to jointly perform KWS and users' own-voice/external speaker detection with a negligible increase in the number of parameters. For experiments, we generate from the Google Speech Commands Dataset a speech corpus emulating hearing aids as a capturing device. Our results show that this multi-task deep residual network is able to achieve a KWS accuracy relative improvement of around 32% with respect to a system that does not deal with external speakers.
Iván López-Espejo, Zheng-Hua Tan, Jesper Jensen 0001
INTERSPEECH3
2019 Deep-learning-based audio-visual speech enhancement in presence of Lombard effect
Daniel Michelsanti, Zheng-Hua Tan, Sigurður Sigurðsson, Jesper Jensen 0001
Speech Commun.4
2019 Asymmetric Coding for Rate-Constrained Noise Reduction in Binaural Hearing Aids
abstract
Binaural hearing aids (HAs) can potentially perform advanced noise reduction algorithms, leading to an improvement over monaural/bilateral HAs. Due to the limited transmission capacities between the HAs and given knowledge of the complete joint noisy signal statistics, the optimal rate-constrained beamforming strategy is known from the literature. However, as these joint statistics are unknown in practice, sub-optimal strategies have been presented. In this paper, we present a unified framework to study the performance of these existing optimal and sub-optimal rate-constrained beamforming methods for binaural HAs. Moreover, we propose to use an asymmetric sequential coding scheme to estimate the joint statistics between the microphones in the two HAs. We show that under certain assumptions, this leads to sub-optimal performance in one HA but allows to obtain the truly optimal performance in the second HA. Based on the mean square error distortion measure, we evaluate the performance improvement between monaural beamforming (no communication) and the proposed scheme, as well as the optimal and the existing sub-optimal strategies in terms of the information bit-rate. The results show that the proposed method outperforms existing practical approaches in most scenarios, especially at middle rates and high rates, without having the prior knowledge of the joint statistics.
Jamal Amini, Richard C. Hendriks, Richard Heusdens, Meng Guo 0001, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2019 Information Loss in the Human Auditory System
abstract
From the eardrum to the auditory cortex, where acoustic stimuli are decoded, there are several stages of auditory processing and transmission where information may potentially be lost. In this paper, we aim at quantifying the total information loss in the human auditory system by using information theoretic tools. To do so, we consider a speech communication model, where words are uttered and sent through a noisy channel, and then received and processed by a human listener. We define a notion of information loss that is related to the human word recognition rate. To assess the word recognition rate of humans, we conduct a closed-vocabulary intelligibility test. We derive upper and lower bounds on the information loss. Simulations reveal that the bounds are tight and we observe that the information loss in the human auditory system increases as the signal to noise ratio (SNR) decreases. Our framework also allows us to study whether humans are optimal in terms of speech perception in a noisy environment. Toward that end, we derive optimal classifiers and compare the human and machine performance in terms of information loss and word recognition rate. We observe a higher information loss and lower word recognition rate for humans compared to the optimal classifiers. In fact, depending on the SNR, the machine classifier may outperform humans by as much as 8 dB. This implies that for the speech-in-stationary-noise setup considered here, the human auditory system is suboptimal for recognizing noisy words.
Mohsen Zareian Jahromi, Adel Zahedi, Jesper Jensen 0001, Jan Østergaard
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 On the Relationship Between Short-Time Objective Intelligibility and Short-Time Spectral-Amplitude Mean-Square Error for Speech Enhancement
abstract
The majority of deep neural network (DNN) based speech enhancement algorithms rely on the mean-square error (MSE) criterion of short-time spectral amplitudes (STSA), which has no apparent link to human perception, e.g., speech intelligibility. Short-time objective intelligibility (STOI), a popular state-of-the-art speech intelligibility estimator, on the other hand, relies on linear correlation of speech temporal envelopes. This raises the question if a DNN training criterion based on envelope linear correlation (ELC) can lead to improved speech intelligibility performance of DNN-based speech enhancement algorithms compared to algorithms based on the STSA-MSE criterion. In this paper, we derive that, under certain general conditions, the STSA-MSE and ELC criteria are practically equivalent, and we provide empirical data to support our theoretical results. Furthermore, our experimental findings suggest that the standard STSA minimum-MSE estimator is near optimal, if the objective is to enhance noisy speech in a manner, which is optimal with respect to the STOI speech intelligibility estimator.
Morten Kolbæk, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 A Convex Approximation of the Relaxed Binaural Beamforming Optimization Problem
abstract
The recently proposed relaxed binaural beamforming (RBB) optimization problem provides a flexible tradeoff between noise suppression and binaural-cue preservation of the sound sources in the acoustic scene. It minimizes the output noise power, under the constraints, which guarantee that the target remains unchanged after processing and the binaural-cue distortions of the acoustic sources will be less than a user-defined threshold. However, the RBB problem is a computationally demanding non convex optimization problem. The only existing suboptimal method which approximately solves the RBB is a successive convex optimization (SCO) method which, typically, requires to solve multiple convex optimization problems per frequency bin, in order to converge. Convergence is achieved when all constraints of the RBB optimization problem are satisfied. In this paper, we propose a semidefinite convex relaxation (SDCR) of the RBB optimization problem. The proposed suboptimal SDCR method solves a single convex optimization problem per frequency bin, resulting in a much lower computational complexity than the SCO method. Unlike the SCO method, the SDCR method does not guarantee user-controlled upper-bounded binaural-cue distortions. To tackle this problem, we also propose a suboptimal hybrid method that combines the SDCR and SCO methods. Instrumental measures combined with a listening test show that the SDCR and hybrid methods achieve significantly lower computational complexity than the SCO method, and in most cases better tradeoff between predicted intelligibility and binaural-cue preservation than the SCO method.
Andreas I. Koutrouvelis, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Robust Joint Estimation of Multimicrophone Signal Model Parameters
abstract
One of the biggest challenges in multimicrophone applications is the estimation of the parameters of the signal model, such as the power spectral densities (PSDs) of the sources, the early (relative) acoustic transfer functions of the sources with respect to the microphones, the PSD of late reverberation, and the PSDs of microphone-self noise. Typically, existing methods estimate subsets of the aforementioned parameters and assume some of the other parameters to be known a priori. This may result in inconsistencies and inaccurately estimated parameters and potential performance degradation in the applications using these estimated parameters. So far, there is no method to jointly estimate all the aforementioned parameters. In this paper, we propose a robust method for jointly estimating all the aforementioned parameters using confirmatory factor analysis. The estimation accuracy of the signal-model parameters thus obtained outperforms existing methods in most cases. We experimentally show significant performance gains in several multimicrophone applications over state-of-the-art methods.
Andreas I. Koutrouvelis, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 Monaural Speech Enhancement Using Deep Neural Networks by Maximizing a Short-Time Objective Intelligibility Measure
abstract
In this paper we propose a Deep Neural Network (D NN) based Speech Enhancement (SE) system that is designed to maximize an approximation of the Short-Time Objective Intelligibility (STOI) measure. We formalize an approximate-STOI cost function and derive analytical expressions for the gradients required for DNN training and show that these gradients have desirable properties when used together with gradient based optimization techniques. We show through simulation experiments that the proposed SE system achieves large improvements in estimated speech intelligibility, when tested on matched and unmatched natural noise types, at multiple signal-to-noise ratios. Furthermore, we show that the SE system, when trained using an approximate-STOI cost function performs on par with a system trained with a mean square error cost applied to short-time temporal envelopes. Finally, we show that the proposed SE system performs on par with a traditional DNN based Short- Time Spectral Amplitude (STSA) SE system in terms of estimated speech intelligibility. These results are important because they suggest that traditional DNN based STSA SE systems might be optimal in terms of estimated speech intelligibility.
Morten Kolbæk, Zheng-Hua Tan, Jesper Jensen 0001
ICASSP3
2018 Refinement and validation of the binaural short time objective intelligibility measure for spatially diverse conditions
Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001
Speech Commun.4
2018 Nonintrusive Speech Intelligibility Prediction Using Convolutional Neural Networks
abstract
Speech Intelligibility Prediction (SIP) algorithms are becoming popular tools within the development and operation of speech processing devices and algorithms. However, many SIP algorithms require knowledge of the underlying clean speech; a signal that is often not available in real-world applications. This has led to increased interest in nonintrusive SIP algorithms, which do not require clean speech to make predictions. In this paper, we investigate the use of Convolutional Neural Networks (CNNs) for nonintrusive SIP. To do so, we utilize a CNN architecture that shows similarities to existing SIP algorithms, in terms of computational structure, and which allows for easy and meaningful visualization and interpretation of trained weights. We evaluate this architecture using a large dataset obtained by combining datasets from the literature. The proposed method shows high prediction performance when compared with four existing intrusive and nonintrusive SIP algorithms. This demonstrates the potential of deep learning for speech intelligibility prediction.
Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 Evaluation and Comparison of Late Reverberation Power Spectral Density Estimators
abstract
Reduction of late reverberation can be achieved using spatio-spectral filters, such as the multichannel Wiener filter. To compute this filter, an estimate of the late reverberation power spectral density (PSD) is required. In recent years, a multitude of late reverberation PSD estimators have been proposed. In this paper, these estimators are categorized into several classes, their relations and differences are discussed, and a comprehensive experimental comparison is provided. To compare their performance, simulations in controlled as well as practical scenarios are conducted. It is shown that a common weakness of spatial coherence-based estimators is their performance in high direct-to-diffuse ratio conditions. To mitigate this problem, a correction method is proposed and evaluated. It is shown that the proposed correction method can decrease the speech distortion without significantly affecting the reverberation reduction.
Sebastian Braun, Adam Kuklasinski, Ofer Schwartz, Oliver Thiergart, Emanuël A. P. Habets, Sharon Gannot, Simon Doclo, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.8
2018 Bias-Compensated Informed Sound Source Localization Using Relative Transfer Functions
abstract
In this paper, we consider the problem of estimating the target sound direction of arrival (DoA) for a hearing aid (HA) system, which can connect to a wireless microphone worn by the talker of interest. The wireless microphone “informs” the HA system about the noise-free target speech. To estimate the DoA, we consider a maximum-likelihood approach, and we assume that a database of DoA-dependent relative transfer functions (RTFs) has been measured in advance and is available. The proposed DoA estimator is able to take the available noise-free target speech, ambient noise characteristics, and the shadowing effect of the user's head on the received signals into account, and it supports both monaural and binaural microphone array configurations. Moreover, we analytically analyze the bias in the proposed estimator and introduce a modified estimator, which has been compensated for the bias. We demonstrate that the proposed method has lower computational complexity and better performance than recent RTF-based estimators. Furthermore, to decrease the number of parameters required to be wirelessly exchanged between the HAs in binaural configurations, we propose an information fusion strategy, which avoids transmitting microphone signals between the HAs. An important benefit of the proposed IF strategy is that the number of parameters to be exchanged between the HAs is independent of the number of HA microphones. Finally, we investigate the performance of variants of the proposed estimator extensively in different noisy and reverberant situations.
Mojtaba Farmani, Michael Syskind Pedersen, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 A non-intrusive Short-Time Objective Intelligibility measure
abstract
We propose a non-intrusive intelligibility measure for noisy and non-linearly processed speech, i.e. a measure which can predict intelligibility from a degraded speech signal without requiring a clean reference signal. The proposed measure is based on the Short-Time Objective Intelligibility (STOI) measure. In particular, the non-intrusive STOI measure estimates clean signal amplitude envelopes from the degraded signal. Subsequently, the STOI measure is evaluated by use of the envelopes of the degraded signal and the estimated clean envelopes. The performance of the proposed measure is evaluated on a dataset including speech in different noise types, processed with binary masks. The measure is shown to predict intelligibility well in all tested conditions, with the exception of those including a single competing speaker. While the measure does not perform as well as the original (intrusive) STOI measure, it is shown to outperform existing non-intrusive measures.
Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001
ICASSP4
2017 Permutation invariant training of deep models for speaker-independent multi-talker speech separation
abstract
We propose a novel deep learning training criterion, named permutation invariant training (PIT), for speaker independent multi-talker speech separation, commonly known as the cocktail-party problem. Different from the multi-class regression technique and the deep clustering (DPCL) technique, our novel approach minimizes the separation error directly. This strategy effectively solves the long-lasting label permutation problem, that has prevented progress on deep learning based techniques for speech separation. We evaluated PIT on the WSJ0 and Danish mixed-speech separation tasks and found that it compares favorably to non-negative matrix factorization (NMF), computational auditory scene analysis (CASA), and DPCL and generalizes well over unseen speakers and languages. Since PIT is simple to implement and can be easily integrated and combined with other advanced techniques, we believe improvements built upon PIT can eventually solve the cocktail-party problem.
Dong Yu 0001, Morten Kolbæk, Zheng-Hua Tan, Jesper Jensen 0001
ICASSP4
2017 On the Use of Band Importance Weighting in the Short-Time Objective Intelligibility Measure
abstract
Speech intelligibility prediction methods are popular tools within the speech processing community for objective evaluation of speech intelligibility of e.g.enhanced speech.The Short-Time Objective Intelligibility (STOI) measure has become highly used due to its simplicity and high prediction accuracy.In this paper we investigate the use of Band Importance Functions (BIFs) in the STOI measure, i.e. of unequally weighting the contribution of speech information from each frequency band.We do so by fitting BIFs to several datasets of measured intelligibility, and cross evaluating the prediction performance.Our findings indicate that it is possible to improve prediction performance in specific situations.However, it has not been possible to find BIFs which systematically improve prediction performance beyond the data used for fitting.In other words, we find no evidence that the performance of the STOI measure can be improved considerably by extending it with a non-uniform BIF.
Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001
INTERSPEECH4
2017 Humans do not Maximize the Probability of Correct Decision When Recognizing DANTALE Words in Noise
abstract
Inspired by the DANTALE II listening test paradigm, which is used for determining the intelligibility of noisy speech, we assess the hypothesis that humans maximize the probability of correct decision when recognizing words contaminated by additive Gaussian, speech-shaped noise.We first propose a statistical Gaussian communication and classification scenario, where word models are built from short term spectra of human speech, and optimal classifiers in the sense of maximum a posteriori estimation are derived.Then, we perform a listening test, where the participants are instructed to make their best guess of words contaminated with speech-shaped Gaussian noise.Comparing the human's performance to that of the optimal classifier reveals that at high SNR, humans perform comparable to the optimal classifier.However, at low SNR, the human performance is inferior to that of the optimal classifier.This shows that, at least in this specialized task, humans are generally not able to maximize the probability of correct decision, when recognizing words.
Mohsen Zareian Jahromi, Jan Østergaard, Jesper Jensen 0001
INTERSPEECH3
2017 Informed Sound Source Localization Using Relative Transfer Functions for Hearing Aid Applications
abstract
Recent hearing aid systems (HASs) can connect to a wireless microphone worn by the talker of interest. This feature gives the HASs access to a noise-free version of the target signal. In this paper, we address the problem of estimating the target sound direction of arrival (DoA) for a binaural HAS given access to the noise-free content of the target signal. To estimate the DoA, we present a maximum-likelihood framework which takes the shadowing effect of the user's head on the received signals into account by modeling the relative transfer functions (RTFs) between the HAS's microphones. We propose three different RTF models which have different degrees of accuracy and individualization. Furthermore, we show that the proposed DoA estimators can be formulated in terms of inverse discrete Fourier transforms to evaluate the likelihood function computationally efficiently. We extensively assess the performance of the proposed DoA estimators for various DoAs, signal to noise ratios, and in different noisy and reverberant situations. The results show that the proposed estimators improve the performance markedly over other recently proposed “informed” DoA estimator.
Mojtaba Farmani, Michael Syskind Pedersen, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Speech Intelligibility Potential of General and Specialized Deep Neural Network Based Speech Enhancement Systems
abstract
In this paper, we study aspects of single microphone speech enhancement (SE) based on deep neural networks (DNNs). Specifically, we explore the generalizability capabilities of state-of-the-art DNN-based SE systems with respect to the background noise type, the gender of the target speaker, and the signal-to-noise ratio (SNR). Furthermore, we investigate how specialized DNN-based SE systems, which have been trained to be either noise type specific, speaker specific or SNR specific, perform relative to DNN based SE systems that have been trained to be noise type general, speaker general, and SNR general. Finally, we compare how a DNN-based SE system trained to be noise type general, speaker general, and SNR general performs relative to a state-of-the-art short-time spectral amplitude minimum mean square error (STSA-MMSE) based SE algorithm. We show that DNN-based SE systems, when trained specifically to handle certain speakers, noise types and SNRs, are capable of achieving large improvements in estimated speech quality (SQ) and speech intelligibility (SI), when tested in matched conditions. Furthermore, we show that improvements in estimated SQ and SI can be achieved by a DNN-based SE system when exposed to unseen speakers, genders and noise types, given a large number of speakers and noise types have been used in the training of the system. In addition, we show that a DNN-based SE system that has been trained using a large number of speakers and a wide range of noise types outperforms a state-of-the-art STSA-MMSE based SE method, when tested using a range of unseen speakers and noise types. Finally, a listening test using several DNN-based SE systems tested in unseen speaker conditions show that these systems can improve SI for some SNR and noise type configurations but degrade SI for others.
Morten Kolbæk, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Multitalker Speech Separation With Utterance-Level Permutation Invariant Training of Deep Recurrent Neural Networks
abstract
In this paper, we propose the utterance-level permutation invariant training (uPIT) technique. uPIT is a practically applicable, end-to-end, deep-learning-based solution for speaker independent multitalker speech separation. Specifically, uPIT extends the recently proposed permutation invariant training (PIT) technique with an utterance-level cost function, hence eliminating the need for solving an additional permutation problem during inference, which is otherwise required by frame-level PIT. We achieve this using recurrent neural networks (RNNs) that, during training, minimize the utterance-level separation error, hence forcing separated frames belonging to the same speaker to be aligned to the same output stream. In practice, this allows RNNs, trained with uPIT, to separate multitalker mixed speech without any prior knowledge of signal duration, number of speakers, speaker identity, or gender. We evaluated uPIT on the WSJ0 and Danish two- and three-talker mixed-speech separation tasks and found that uPIT outperforms techniques based on nonnegative matrix factorization and computational auditory scene analysis, and compares favorably with deep clustering, and the deep attractor network. Furthermore, we found that models trained with uPIT generalize well to unseen speakers and languages. Finally, we found that a single model, trained with uPIT, can handle both two-speaker, and three-speaker speech mixtures.
Morten Kolbæk, Dong Yu 0001, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Relaxed Binaural LCMV Beamforming
abstract
In this paper, we propose a new binaural beamforming technique, which can be seen as a relaxation of the linearly constrained minimum variance (LCMV) framework. The proposed method can achieve simultaneous noise reduction and exact binaural cue preservation of the target source, similar to the binaural minimum variance distortionless response (BMVDR) method. However, unlike BMVDR, the proposed method is also able to preserve the binaural cues of multiple interferers to a certain predefined accuracy. Specifically, it is able to control the trade-off between noise reduction and binaural cue preservation of the interferers by using a separate trade-off parameter per-interferer. Moreover, we provide a robust way of selecting these trade-off parameters in such a way that the preservation accuracy for the binaural cues of the interferers is always better than the corresponding ones of the BMVDR. The relaxation of the constraints in the proposed method achieves approximate binaural cue preservation of more interferers than other previously presented LCMV-based binaural beamforming methods that use strict equality constraints.
Andreas I. Koutrouvelis, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 Concurrent localization of sound sources and dual-microphone sub-arrays using TOFs
Mojtaba Farmani, Richard Heusdens, Michael Syskind Pedersen, Zheng-Hua Tan, Jesper Jensen 0001
FUSION5
2016 A method for predicting the intelligibility of noisy and non-linearly enhanced binaural speech
abstract
We propose and evaluate a binaural speech intelligibility measure. The measure is a binaural extension of the Short-Time Objective Intelligibility (STOI) measure and focuses on predicting the intelligibility of noisy speech which has been enhanced by a speech processing algorithm (e.g. in a hearing aid). We show that the measure can accurately predict 1) the Speech Reception Threshold (SRT) for a frontal speaker masked by a point noise source in the horizontal plane, 2) the improvement in SRT obtained by independently processing the left and right ear signals with Ideal Time Frequency Segregation (ITFS), and 3) the intelligibility of speech in the presence of multiple interferers as well as the effect of processing the noisy signals with 2-microphone MVDR beamforming as used in hearing aids. Finally, we show that the computational demands associated with the measure are favourable in comparison with those of a previously proposed measure with similar properties.
Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001
ICASSP4
2016 Informed Direction of Arrival estimation using a spherical-head model for Hearing Aid applications
abstract
In this paper, we propose a Direction of Arrival (DoA) estimator for a Hearing Aid System (HAS) which can connect to a wireless microphone worn by a target talker. The wireless microphone "informs" the HAS about the almost noise-free content of the target sound, and the proposed DoA estimator uses the knowledge of the noise-free target sound and the received microphone signals to estimate the DoA via a maximum likelihood approach. Moreover, the proposed DoA estimator resorts to a user-independent spherical-head model to consider the acoustic impacts of the head on the received signals at the HAS. Further, the proposed DoA estimator uses an Inverse Discrete Fourier Transform (IDFT) technique to evaluate the likelihood function computationally efficiently. We assessed the performance of the proposed estimator for various DoAs, Signal to Noise Ratios (SNRs), and target distances in different noisy and reverberant situations. The proposed estimator improves the performance markedly over other recently proposed "informed" DoA estimators.
Mojtaba Farmani, Michael Syskind Pedersen, Zheng-Hua Tan, Jesper Jensen 0001
ICASSP4
2016 Improved multi-microphone noise reduction preserving binaural cues
abstract
We propose a new multi-microphone noise reduction technique for binaural cue preservation of the desired source and the interferers. This method is based on the linearly constrained minimum variance (LCMV) framework, where the constraints are used for the binaural cue preservation of the desired source and of multiple interferers. In this framework there is a trade-off between noise reduction and binaural cue preservation. The more constraints the LCMV uses for preserving binaural cues, the less degrees of freedom can be used for noise suppression. The recently presented binaural LCMV (BLCMV) method and the optimal BLCMV (OBLCMV) method require two constraints per interferer and introduce an additional interference rejection parameter. This unnecessarily reduces the degrees of freedom, available for noise reduction, and negatively influences the trade-off between noise reduction and binaural cue preservation. With the proposed method, binaural cue preservation is obtained using just a single constraint per interferer without the need of an interference rejection parameter. The proposed method can simultaneously achieve noise reduction and perfect binaural cue preservation of more than twice as many interferers as the BLCMV, while the OBLCMV can preserve the binaural cues of only one interferer.
Andreas I. Koutrouvelis, Richard C. Hendriks, Jesper Jensen 0001, Richard Heusdens
ICASSP3
2016 Maximum likelihood PSD estimation for speech enhancement in reverberant and noisy conditions
abstract
We propose a novel Power Spectral Density (PSD) estimator for multi-microphone systems operating in reverberant and noisy conditions. The estimator is derived using the maximum likelihood approach and is based on a blocked and pre-whitened additive signal model. The intended application of the estimator is in speech enhancement algorithms, such as the Multi-channel Wiener Filter (MWF) and the Minimum Variance Distortionless Response (MVDR) beamformer. We evaluate these two algorithms in a speech dereverberation task and compare the performance obtained using the proposed and a competing PSD estimator. Instrumental performance measures indicate an advantage of the proposed estimator over the competing one. In a speech intelligibility test all algorithms significantly improved the word intelligibility score. While the results suggest a minor advantage of using the proposed PSD estimator, the difference between algorithms was found to be statistically significant only in some of the experimental conditions.
Adam Kuklasinski, Simon Doclo, Jesper Jensen 0001
ICASSP3
2016 Speech enhancement using Long Short-Term Memory based recurrent Neural Networks for noise robust Speaker Verification
abstract
In this paper we propose to use a state-of-the-art Deep Recurrent Neural Network (DRNN) based Speech Enhancement (SE) algorithm for noise robust Speaker Verification (SV). Specifically, we study the performance of an i-vector based SV system, when tested in noisy conditions using a DRNN based SE front-end utilizing a Long Short-Term Memory (LSTM) architecture. We make comparisons to systems using a Non-negative Matrix Factorization (NMF) based front-end, and a Short-Time Spectral Amplitude Minimum Mean Square Error (STSA-MMSE) based front-end, respectively. We show in simulation experiments that a male-speaker and text-independent DRNN based SE front-end, without specific a priori knowledge about the noise type outperforms a text, noise type and speaker dependent NMF based front-end as well as a STSA-MMSE based front-end in terms of Equal Error Rates for a large range of noise types and signal to noise ratios on the RSR2015 speech corpus.
Morten Kolbæk, Zheng-Hua Tan, Jesper Jensen 0001
SLT3
2016 Predicting the Intelligibility of Noisy and Nonlinearly Processed Binaural Speech
abstract
Objective speech intelligibility measures are gaining popularity in the development of speech enhancement algorithms and speech processing devices such as hearing aids. Such devices may process the input signals nonlinearly and modify the binaural cues presented to the user. We propose a method for predicting the intelligibility of noisy and nonlinearly processed binaural speech. This prediction is based on the noisy and processed signal as well as a clean speech reference signal. The method is obtained by extending a modified version of the short-time objective intelligibility (STOI) measure with a modified equalization-cancellation (EC) stage. We evaluate the performance of the method by comparing the predictions with measured intelligibility from four listening experiments. These comparisons indicate that the proposed measure can provide accurate predictions of (1) the intelligibility of diotic speech with an accuracy similar to that of the original STOI measure, (2) speech reception thresholds (SRTs) in conditions with a frontal target speaker and a single interferer in the horizontal plane, (3) SRTs in conditions with a frontal target and a single interferer when ideal time frequency segregation (ITFS) is applied to the left and right ears separately, and (4) the advantage of two-microphone beamforming as applied in state-of-the-art hearing aids. A MATLAB implementation of the proposed measure is available online1.
Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 An Algorithm for Predicting the Intelligibility of Speech Masked by Modulated Noise Maskers
abstract
Intelligibility listening tests are necessary during development and evaluation of speech processing algorithms, despite the fact that they are expensive and time consuming. In this paper, we propose a monaural intelligibility prediction algorithm, which has the potential of replacing some of these listening tests. The proposed algorithm shows similarities to the short-time objective intelligibility (STOI) algorithm, but works for a larger range of input signals. In contrast to STOI, extended STOI (ESTOI) does not assume mutual independence between frequency bands. ESTOI also incorporates spectral correlation by comparing complete 400ms length spectrograms of the noisy/processed speech and the clean speech signals. As a consequence, ESTOI is also able to accurately predict the intelligibility of speech contaminated by temporally highly modulated noise sources in addition to noisy signals processed with time-frequency weighting. We show that ESTOI can be interpreted in terms of an orthogonal decomposition of short-time spectrograms into intelligibility subspaces, i.e., a ranking of spectrogram features according to their importance to intelligibility. A free MATLAB implementation of the algorithm is available for noncommercial use at http://kom.aau.dk/~jje/.
Jesper Jensen 0001, Cees H. Taal
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Maximum Likelihood PSD Estimation for Speech Enhancement in Reverberation and Noise
abstract
In this contribution, we focus on the problem of power spectral density (PSD) estimation from multiple microphone signals in reverberant and noisy environments. The PSD estimation method proposed in this paper is based on the maximum likelihood (ML) methodology. In particular, we derive a novel ML PSD estimation scheme that is suitable for sound scenes which besides speech and reverberation consists of an additional noise component whose second-order statistics are known. The proposed algorithm is shown to outperform an existing similar algorithm in terms of PSD estimation accuracy. Moreover, it is shown numerically that the mean-squared estimation error achieved by the proposed method is near the limit set by the corresponding Cramér-Rao lower bound. The speech dereverberation performance of a multichannel Wiener filter based on the proposed PSD estimators is measured using several instrumental measures and is shown to be higher than when the competing estimator is used. Moreover, we perform a speech intelligibility test where we demonstrate that both the proposed and the competing PSD estimators lead to similar intelligibility improvements.
Adam Kuklasinski, Simon Doclo, Søren Holdt Jensen, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2015 Maximum likelihood approach to "informed" Sound Source Localization for Hearing Aid applications
abstract
Most state-of-the-art Sound Source Localization (SSL) algorithms have been proposed for applications which are “uninformed” about the target sound content; however, utilizing a wireless microphone worn by a target talker, enables recent Hearing Aid Systems (HASs) to access to an almost noise-free sound signal of the target talker at the HAS via the wireless connection. Therefore, in this paper, we propose a maximum likelihood (ML) approach, which we call MLSSL, to estimate the Direction of Arrival (DoA) of the target signal given access to the target signal content. Compared with other “informed” SSL algorithms which use binaural microphones for localization, MLSSL performs better using signals of one or more microphones placed on just one ear, thereby reducing the wireless transmission overhead of binaural hearing aids. More specifically, when the target location confined to the front-horizontal plane, MLSSL shows an average absolute DoA estimation error of 5 degrees at SNR of -5 dB in a large-crowd noise and non-reverberant situation. Moreover, MLSSL suffers less from front-back confusions compared with the recent approaches.
Mojtaba Farmani, Michael Syskind Pedersen, Zheng-Hua Tan, Jesper Jensen 0001
ICASSP4
2015 On the influence of microphone array geometry on HRTF-based Sound Source Localization
abstract
The direction dependence of Head Related Transfer Functions (HRTFs) forms the basis for HRTF-based Sound Source Localization (SSL) algorithms. In this paper, we show how spectral similarities of the HRTFs of different directions in the horizontal plane influence performance of HRTF-based SSL algorithms; the more similar the HRTFs of different angles to the HRTF of the target angle, the worse the performance. However, we also show how the microphone array geometry can assist in differentiating between the HRTFs of the different angles, thereby improving performance of HRTF-based SSL algorithms. Furthermore, to demonstrate the analysis results, we show the impact of HRTFs similarities and microphone array geometry on an exemplary HRTF-based SSL algorithm, called MLSSL. This algorithm is well-suited for this purpose as it allows to estimate the Direction-of-Arrival (DoA) of the target sound using any number of microphones and any geometries of the microphone array around the head.
Mojtaba Farmani, Michael Syskind Pedersen, Zheng-Hua Tan, Jesper Jensen 0001
ICASSP4
2015 A simple modification to facilitate robust generalized sidelobe canceller for hearing aids
abstract
This work focuses on an adaptive beamformer in a hearing aid application using a generalized sidelobe canceller structure (GSC). In this application, the constraint and blocking matrices in the GSC structure are specifically designed using an estimate of the transfer functions between the target source and the microphones to ensure optimal beamformer performance. We show that, in practice, the GSC always-unintentionally-attenuates the target sound in a special but realistic situation where all signals, including the target and noise signals, originate from the look direction reflected by the look vector. This happens because, in practice, the blocking matrix in the GSC structure is non-ideal. We introduce a simple modification to the GSC structure, which solves the problem of undesired target signal attenuation in situations where all signals originate from the look direction. Furthermore, this modification can also prevent desired signals, originating from positions spatially close to the look direction, to be removed. We also show that the solution has no impact on other acoustic situations.
Meng Guo 0001, Jan Mark de Haan, Jesper Jensen 0001
ICASSP3
2015 Speech reinforcement in noisy reverberant conditions under an approximation of the short-time SII
abstract
While most contributions on speech reinforcement only consider the presence of environmental noise, late reverberation can also severely degrade the intelligibility of speech. In this paper we address the problem of speech reinforcement in noisy and reverberant environments. We use a short-time version of a recently presented approximation of the speech intelligibility index, which we optimize locally. The resulting time-frequency dependent amplification depends on both the noise and late reverberation power spectral density. The latter is estimated using the Polack model and assumes that prior knowledge of the room geometry is available. Speech intelligibility improvements of around 20% are observed.
Richard C. Hendriks, Joao B. Crespo, Jesper Jensen 0001, Cees H. Taal
ICASSP3
2015 Analysis of beamformer directed single-channel noise reduction system for hearing aid applications
abstract
We study multi-microphone noise reduction systems consisting of a beamformer and a single-channel (SC) noise reduction stage. In particular, we present and analyse a maximum likelihood (ML) method for jointly estimating the target and noise power spectral densities (psd's) entering the SC filter. We show that the estimators are minimum variance and unbiased, and provide closed-form expressions for their mean-square error (MSE). Furthermore, we show that the MSE of the noise psd estimator is particularly simple: it is independent of target signal characteristics, frequency, and microphone locations. In a hearing aid context, we analyze the performance of the estimators as a function of target angle-of-arrival and frequency. Finally, we demonstrate the advantage of the proposed method in a hearing aid situation with a target speaker in large-crowd noise.
Jesper Jensen 0001, Michael Syskind Pedersen
ICASSP1
2015 Multi-channel PSD estimators for speech dereverberation - A theoretical and experimental comparison
abstract
In this paper we perform an extensive theoretical and experimental comparison of two recently proposed multi-channel speech dereverberation algorithms. Both of them are based on the multi-channel Wiener filter but they use different estimators of the speech and reverberation power spectral densities (PSDs). We first derive closedform expressions for the mean square error (MSE) of both PSD estimators and then show that one estimator - previously used for speech dereverberation by the authors - always yields a better MSE. Only in the case of a two microphone array or for special spatial distributions of the interference both estimators yield the same MSE. The theoretically derived MSE values are in good agreement with numerical simulation results and with instrumental speech quality measures in a realistic speech dereverberation task for binaural hearing aids.
Adam Kuklasinski, Simon Doclo, Timo Gerkmann, Søren Holdt Jensen, Jesper Jensen 0001
ICASSP5
2015 A binaural short time objective intelligibility measure for noisy and enhanced speech
abstract
Objective intelligibility measures are increasingly being used to assess the performance of speech processing algorithms, e.g. for hearing aids. It has been shown that the short time objective intelligibility (STOI) measure yields good results in this respect. In this paper we propose a binaural extension of the STOI measure, which predicts binaural advantage using a modified equalization cancellation (EC) stage. The proposed method is evaluated for a range of acoustic conditions. Firstly, the method is able to predict the advantage of spatial separation between a speech target and a speech shaped noise (SSN) interferer. Secondly, the method yields results comparable to the monaural STOI measure when presented with noisy speech processed by ideal time-frequency segregation (ITFS). Finally, the method also performs well when presented with a selection of different acoustic conditions combined with beamforming as used in hearing aids.
Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, Jesper Jensen 0001
INTERSPEECH4
2015 Optimal Near-End Speech Intelligibility Improvement Incorporating Additive Noise and Late Reverberation Under an Approximation of the Short-Time SII
abstract
The presence of environmental additive noise in the vicinity of the user typically degrades the speech intelligibility of speech processing applications. This intelligibility loss can be compensated by properly preprocessing the speech signal prior to play-out, often referred to as near-end speech enhancement. Although the majority of such algorithms focus primarily on the presence of additive noise, reverberation can also severely degrade intelligibility. In this paper we investigate how late reverberation and additive noise can be jointly taken into account in the near-end speech enhancement process. For this effort we use a recently presented approximation of the speech intelligibility index under a power constraint, which we optimize for speech degraded by both additive noise and late reverberation. The algorithm results in time-frequency dependent amplification factors that depend on both the additive noise power spectral density as well as the late reverberation energy. These amplification factors redistribute speech energy across frequency and perform a dynamic range compression. Experimental results using both instrumental intelligibility measures as well as intelligibility listening tests show that the proposed approach improves speech intelligibility over state-of-the-art reference methods when speech signals are degraded simultaneously by additive noise and reverberation. Speech intelligibility improvements in the order of 20% are observed.
Richard C. Hendriks, Joao B. Crespo, Jesper Jensen 0001, Cees H. Taal
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Minimum Mean-Square Error Estimation of Mel-Frequency Cepstral Features-A Theoretically Consistent Approach
abstract
In this work, we consider the problem of feature enhancement for noise-robust automatic speech recognition (ASR). We propose a method for minimum mean-square error (MMSE) estimation of mel-frequency cepstral features, which is based on a minimum number of well-established, theoretically consistent statistical assumptions. More specifically, the method belongs to the class of methods relying on the statistical framework proposed in Ephraim and Malah's original work (“Speech enhancement using a minimum mean-square error short-time spectral amplitude estimator,” IEEE Trans. Acoust., Speech, Signal Process., vol. ASSP-32, no. 6, 1984). The method is general in that it allows MMSE estimation of mel-frequency cepstral coefficients (MFCC's), cepstral-mean subtracted (CMS-) MFCC's, autoregressive-moving-average (ARMA)-filtered CMS-MFCC's, velocity, and acceleration coefficients. In addition, the method is easily modified to take into account other compressive non-linearities than the logarithm traditionally used for MFCC computation. In terms of MFCC estimation performance, as measured by MFCC mean-square error, the proposed method shows performance which is identical to or better than other state-of-the-art methods. In terms of ASR performance, no statistical difference could be found between the proposed method and the state-of-the-art methods. We conclude that existing state-of-the-art MFCC feature enhancement algorithms within this class of algorithms, while theoretically suboptimal or based on theoretically inconsistent assumptions, perform close to optimally in the MMSE sense.
Jesper Jensen 0001, Zheng-Hua Tan
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Speech Intelligibility Prediction Based on Mutual Information
abstract
This paper deals with the problem of predicting the average intelligibility of noisy and potentially processed speech signals, as observed by a group of normal hearing listeners. We propose a model which performs this prediction based on the hypothesis that intelligibility is monotonically related to the mutual information between critical-band amplitude envelopes of the clean signal and the corresponding noisy/processed signal. The resulting intelligibility predictor turns out to be a simple function of the mean-square error (mse) that arises when estimating a clean critical-band amplitude using a minimum mean-square error (mmse) estimator based on the noisy/processed amplitude. The proposed model predicts that speech intelligibility cannot be improved by any processing of noisy critical-band amplitudes. Furthermore, the proposed intelligibility predictor performs well ( ρ > 0.95) in predicting the intelligibility of speech signals contaminated by additive noise and potentially non-linearly processed using time-frequency weighting.
Jesper Jensen 0001, Cees H. Taal
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Analysis of closed-loop acoustic feedback cancellation systems
abstract
In a previous study, the performance of an acoustic feedback/echo cancellation system was analyzed using a power transfer function method. Whereas the analysis result provides very accurate performance predictions in open-loop acoustic echo cancellation systems, it is less accurate in closed-loop acoustic feedback cancellation systems if there is a strong correlation between the loudspeaker signal and the signals entering the microphones. This work extends the performance analysis to include the effects of the nonzero correlation on the adaptive filters. Simulation results verify that this extension provides much more accurate performance predictions in closed-loop acoustic feedback cancellation systems.
Meng Guo 0001, Søren Holdt Jensen, Jesper Jensen 0001, Steven L. Grant
ICASSP3
2013 Prediction of intelligibility of noisy and time-frequency weighted speech based on mutual information between amplitude envelopes
abstract
This paper deals with the problem of predicting the average intelligibility of noisy and potentially processed speech signals, as observed by a group of normal hearing listeners. We propose a prediction model based on the hypothesis that intelligibility is monotonically related to the the amount of Shannon information the critical-band amplitude envelopes of the noisy/processed signal convey about the corresponding clean signal envelopes. The resulting intelligibility predictor turns out to be a simple function of the correlation between noisy/processed and clean amplitude envelopes. The proposed predictor performs well (ρ>0.95) in predicting the intelligibility of speech signals contaminated by additive noise and potentially non-linearly processed using time-frequency weighting.
Jesper Jensen 0001, Cees H. Taal
INTERSPEECH1
2013 SII-based speech preprocessing for intelligibility improvement in noise
abstract
A linear time-invariant filter is designed in order to improve speech understanding when the speech is played back in a noisy environment. To accomplish this, the speech intelligibility index (SII) is maximized under the constraint that the speech energy is held constant. A nonlinear approximation is used for the SII such that a closed-form solution exists to the constrained optimization problem. The resulting filter is dependent both on the long-term average noise and speech spectrum and the global SNR and, in general, has a high-pass characteristic. In contrast to existing methods, the proposed filter sets certain frequency bands to zero when they do not contribute to intelligibility anymore. Experiments show large intelligibility improvements with the proposed method when used in stationary speech-shaped noise. However, it was also found that the method does not perform well for speech corrupted by a competing speaker. This is due to the fact that the SII is not a reliable intelligibility predictor for fluctuating noise sources. MATLAB code is provided. Index Terms: Speech intelligibility, speech enhancement, nearend enhancement, speech intelligibility index
Cees H. Taal, Jesper Jensen 0001
INTERSPEECH2
2013 On Optimal Linear Filtering of Speech for Near-End Listening Enhancement
abstract
In this letter the focus is on linear filtering of speech before degradation due to additive background noise. The goal is to design the filter such that the speech intelligibility index (SII) is maximized when the speech is played back in a known noisy environment. Moreover, a power constraint is taken into account to prevent uncomfortable playback levels and deal with loudspeaker constraints. Previous methods use linear approximations of the SII in order to find a closed-form solution. However, as we show, these linear approximations introduce errors in low SNR regions and are therefore suboptimal. In this work we propose a nonlinear approximation of the SII which is accurate for all SNRs. Experiments show large intelligibility improvements with the proposed method over the unprocessed noisy speech and better performance than one state-of-the art method.
Cees H. Taal, Jesper Jensen 0001, Arne Leijon
IEEE Signal Process. Lett.2
2012 On Acoustic Feedback Cancellation Using Probe Noise in Multiple-Microphone and Single-Loudspeaker Systems
abstract
A probe noise signal can be used in an acoustic feedback cancellation system to prevent biased adaptive estimation of acoustic feedback paths. However, practical experiences and simulation results indicate that whenever a low-level and inaudible probe noise signal is used, the convergence rate of the adaptive estimation is significantly decreased when keeping the steady-state error unchanged. The goal of this work is to derive analytic expressions for the system behavior such as convergence rate and steady-state error for a multiple-microphone and single-loudspeaker audio system, where the acoustic feedback cancellation is carried out using a probe noise signal. The derived results show how different system parameters and signal properties affect the cancellation performance, and the results explain theoretically the decreased convergence rate. Understanding this is important for making further improvements in the existing probe noise approach.
Meng Guo 0001, Thomas Bo Elmedyb, Søren Holdt Jensen, Jesper Jensen 0001
IEEE Signal Process. Lett.4
2012 Novel Acoustic Feedback Cancellation Approaches in Hearing Aid Applications Using Probe Noise and Probe Noise Enhancement
abstract
Adaptive filters are widely used in acoustic feedback cancellation systems and have evolved to be state-of-the-art. One major challenge remaining is that the adaptive filter estimates are biased due to the nonzero correlation between the loudspeaker signals and the signals entering the audio system. In many cases, this bias problem causes the cancellation system to fail. The traditional probe noise approach, where a noise signal is added to the loudspeaker signal can, in theory, prevent the bias. However, in practice, the probe noise level must often be so high that the noise is clearly audible and annoying; this makes the traditional probe noise approach less useful in practical applications. In this work, we explain theoretically the decreased convergence rate when using low-level probe noise in the traditional approach, before we propose and study analytically two new probe noise approaches utilizing a combination of specifically designed probe noise signals and probe noise enhancement. Despite using low-level and inaudible probe noise signals, both approaches significantly improve the convergence behavior of the cancellation system compared to the traditional probe noise approach. This makes the proposed approaches much more attractive in practical applications. We demonstrate this through a simulation experiment with audio signals in a hearing aid acoustic feedback cancellation system, where the convergence rate is improved by as much as a factor of 10.
Meng Guo 0001, Søren Holdt Jensen, Jesper Jensen 0001
IEEE Trans. Speech Audio Process.3
2012 Spectral Magnitude Minimum Mean-Square Error Estimation Using Binary and Continuous Gain Functions
abstract
Recently, binary mask techniques have been proposed as a tool for retrieving a target speech signal from a noisy observation. A binary gain function is applied to time-frequency tiles of the noisy observation in order to suppress noise dominated and retain target dominated time-frequency regions. When implemented using discrete Fourier transform (DFT) techniques, the binary mask techniques can be seen as a special case of the broader class of DFT-based speech enhancement algorithms, for which the applied gain function is not constrained to be binary. In this context, we develop and compare binary mask techniques to state-of-the-art continuous gain techniques. We derive spectral magnitude minimum mean-square error binary gain estimators; the binary gain estimators turn out to be simple functions of the continuous gain estimators. We show that the optimal binary estimators are closely related to a range of existing, heuristically developed, binary gain estimators. The derived binary gain estimators perform better than existing binary gain estimators in simulation experiments with speech signals contaminated by several different noise sources as measured by speech quality and intelligibility measures. However, even the best binary mask method is significantly outperformed by state-of-the-art continuous gain estimators. The instrumental intelligibility results are confirmed in an intelligibility listening test.
Jesper Jensen 0001, Richard C. Hendriks
IEEE Trans. Speech Audio Process.1
2011 Analysis of adaptive feedback and echo cancelation algorithms in a general multiple-microphone and single-loudspeaker system
abstract
In this paper, we analyze a general multiple-microphone and single-loudspeaker system, where an adaptive algorithm is used to cancel acoustic feedback/echo and a beamformer processes the feedback/echo canceled signals. This system can be viewed as part of a typical hearing aid system and/or a traditional acoustic echo cancellation system. We introduce and derive an approximation of a useful frequency domain measure - the power transfer function - and show how to predict the system stability bound, convergence rate and the steady-state behavior across time and frequency. Furthermore, we show how the derived expressions can be used to determine e.g. the step size parameter in the adaptive algorithms to achieve a desired system property e.g. convergence rate at a specific frequency.
Meng Guo 0001, Thomas Bo Elmedyb, Søren Holdt Jensen, Jesper Jensen 0001
ICASSP4
2011 Spectral magnitude minimum mean-square error binary masks for DFT based speech enhancement
abstract
Originally, ideal binary mask (idbm) techniques have been used as a tool for studying aspects of the auditory system. More recently, idbm techniques have been adapted to the practical problem of retrieving a target speech signal from a noisy observation. In this practical setting, the biliary mask techniques show similarities with existing DFT based speech enhancement techniques. In this context, we derive single-channel, binary mask estimators which minimize the spectral magnitude mean-square error. We show in simulation experiments with natural speech and noise signals that the proposed estimators perform significantly better than existing binary mask estimators. However, even the best of the proposed estimators is clearly out performed by non-binary estimators, both in terms of speech quality and intelligibility.
Jesper Jensen 0001, Richard C. Hendriks
ICASSP1
2011 An Algorithm for Intelligibility Prediction of Time-Frequency Weighted Noisy Speech
abstract
In the development process of noise-reduction algorithms, an objective machine-driven intelligibility measure which shows high correlation with speech intelligibility is of great interest. Besides reducing time and costs compared to real listening experiments, an objective intelligibility measure could also help provide answers on how to improve the intelligibility of noisy unprocessed speech. In this paper, a short-time objective intelligibility measure (STOI) is presented, which shows high correlation with the intelligibility of noisy and time-frequency weighted noisy speech (e.g., resulting from noise reduction) of three different listening experiments. In general, STOI showed better correlation with speech intelligibility compared to five other reference objective intelligibility models. In contrast to other conventional intelligibility models which tend to rely on global statistics across entire sentences, STOI is based on shorter time segments (386 ms). Experiments indeed show that it is beneficial to take segment lengths of this order into account. In addition, a free Matlab implementation is provided.
Cees H. Taal, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
IEEE Trans. Speech Audio Process.4
2010 MMSE based noise PSD tracking with low complexity
abstract
Most speech enhancement algorithms heavily depend on the noise power spectral density (PSD). Because this quantity is unknown in practice, estimation from the noisy data is necessary. We present a low complexity method for noise PSD estimation. The algorithm is based on a minimum mean-squared error estimator of the noise magnitude-squared DFT coefficients. Compared to minimum statistics based noise tracking, segmental SNR and PESQ are improved for non-stationary noise sources with 1 dB and 0.25 MOS points, respectively. Compared to recently published algorithms, similar good noise tracking performance is obtained, but at a computational complexity that is in the order of a factor 40 lower.
Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
ICASSP3
2010 A short-time objective intelligibility measure for time-frequency weighted noisy speech
abstract
Existing objective speech-intelligibility measures are suitable for several types of degradation, however, it turns out that they are less appropriate for methods where noisy speech is processed by a time-frequency (TF) weighting, e.g., noise reduction and speech separation. In this paper, we present an objective intelligibility measure, which shows high correlation (rho=0.95) with the intelligibility of both noisy, and TF-weighted noisy speech. The proposed method shows significantly better performance than three other, more sophisticated, objective measures. Furthermore, it is based on an intermediate intelligibility measure for short-time (approximately 400 ms) TF-regions, and uses a simple DFT-based TF-decomposition. In addition, a free Matlab implementation is provided.
Cees H. Taal, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
ICASSP4
2010 n -Channel Asymmetric Entropy-Constrained Multiple-Description Lattice Vector Quantization
abstract
This paper is about the design and analysis of an index-assignment (IA)-based multiple-description coding scheme for the n-channel asymmetric case. We use entropy constrained lattice vector quantization and restrict attention to simple reconstruction functions, which are given by the inverse IA function when all descriptions are received or otherwise by a weighted average of the received descriptions. We consider smooth sources with finite differential entropy rate and MSE fidelity criterion. As in previous designs, our construction is based on nested lattices which are combined through a single IA function. The results are exact under high-resolution conditions and asymptotically as the nesting ratios of the lattices approach infinity. For any n, the design is asymptotically optimal within the class of IA-based schemes. Moreover, in the case of two descriptions and finite lattice vector dimensions greater than one, the performance is strictly better than that of existing designs. In the case of three descriptions, we show that in the limit of large lattice vector dimensions, points on the inner bound of Pradhan can be achieved. Furthermore, for three descriptions and finite lattice vector dimensions, we show that the IA-based approach yields, in the symmetric case, a smaller rate loss than the recently proposed source-splitting approach.
Jan Østergaard, Richard Heusdens, Jesper Jensen 0001
IEEE Trans. Inf. Theory3
2009 Fast noise PSD estimation with low complexity
abstract
Although noise PSD estimation is a crucial part of noise reduction algorithms, most noise PSD estimators have problems in tracking non-stationary noise sources. Recently, a noise PSD estimator based on DFT-subspace decompositions was proposed, which improves estimation of the PSD of such noise sources. However, as this approach is based on eigenvalue decompositions per DFT bin, it might be too computationally demanding for low-complexity applications like hearing aids. In this paper we present a method with similar noise tracking performance as the DFT-subspace approach, but with low computational costs. This method is based on computation of high resolution perodiograms, and can estimate the noise PSD when both speech and noise are present in a frequency bin. When combined with a complete noise reduction system, the proposed method can lead to an improvement for non-stationary noise sources of more than 1 dB segmental SNR and 0.3 on a PESQ scale, compared to standard noise tracking methods such as minimum statistics and the quantile based approach, while computational complexity is in the same order of magnitude.
Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001, Ulrik Kjems
ICASSP3
2009 Low delay moving-horizon multiple-description audio coding forwireless hearing aids
abstract
In this work, we construct a novel scheme for efficient perceptual coding of audio for robust communication between encoders and wireless hearing aids. To limit the physical size of the hearing aids and to reduce power consumption and thereby increase the lifetime expectancy of the batteries, the hearing aids are constrained to be of low complexity. We therefore provide an asymmetric strategy where most of the computational load is placed at the encoding side. We make use of multiple-description coding. This combats possible erasures on the wireless link between the encoder and the hearing aids without introducing significant delay. Furthermore, we employ psychoacoustically optimized noise-shaping quantizers based on the moving-horizon principle, which exploits a finite prediction horizon.
Jan Østergaard, Daniel E. Quevedo, Jesper Jensen 0001
ICASSP3
2009 Log-spectral magnitude MMSE estimators under super-Gaussian densities
abstract
Despite the fact that histograms of speech DFT coefficients are super-Gaussian, not much attention has been paid to develop estimators under these super-Gaussian distributions in combi-nation with perceptual meaningful distortion measures. In this paper we present log-spectral magnitude MMSE estimators un-der super-Gaussian densities, resulting in an estimator that is perceptually more meaningful and in line with measured his-tograms of speech DFT coefficients. Compared to state-of-the-art reference methods, the presented estimator leads to an im-provement of the segmental SNR in the order of 0.5 dB up to 1 dB. Moreover, listening tests show that the proposed estima-tor leads to significant improvement for the presented estimator over state-of-the-art methods. Index Terms: speech enhancement, log-spectral magnitude MMSE, super-Gaussian
Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
INTERSPEECH3
2009 An evaluation of objective quality measures for speech intelligibility prediction
abstract
In this research various objective quality measures are evalu-ated in order to predict the intelligibility for a wide range of non-linearly processed speech signals and speech degraded by additive noise. The obtained results are compared with the pre-diction results of a more advanced perceptual-based model pro-posed by Dau et al. and an objective intelligibility measure, namely the coherence speech intelligibility index (cSII). These tests are performed in order to gain more knowledge between the link of speech-quality and speech-intelligibility and may help us to exploit the extensive research done into the field of speech-quality for speech-intelligibility. It is shown that cSII does not necessarily show better performance compared to con-ventional objective (speech)-quality measures. In general, the DAU-model is the only method with reasonable results for all processing conditions. Index Terms: Speech intelligibility prediction, speech quality, objective Measure.
Cees H. Taal, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001, Ulrik Kjems
INTERSPEECH4
2009 On Optimal Multichannel Mean-Squared Error Estimators for Speech Enhancement
abstract
In this letter we present discrete Fourier transform (DFT) domain minimum mean-squared error (MMSE) estimators for multichannel noise reduction. The estimators are derived assuming that the clean speech magnitude DFT coefficients are generalized-Gamma distributed. We show that for Gaussian distributed noise DFT coefficients, the optimal filtering approach consists of a concatenation of a minimum variance distortionless response (MVDR) beamformer followed by well-known single-channel MMSE estimators. The multichannel Wiener filter follows as a special case of the presented MSE estimators and is in general suboptimal. For non-Gaussian distributed noise DFT coefficients the resulting spatial filter is in general nonlinear with respect to the noisy microphone signals and cannot be decomposed into an MVDR beamformer and a post-filter.
Richard C. Hendriks, Richard Heusdens, Ulrik Kjems, Jesper Jensen 0001
IEEE Signal Process. Lett.4
2008 Noise Tracking Using DFT Domain Subspace Decompositions
abstract
All discrete Fourier transform (DFT) domain-based speech enhancement gain functions rely on knowledge of the noise power spectral density (PSD). Since the noise PSD is unknown in advance, estimation from the noisy speech signal is necessary. An overestimation of the noise PSD will lead to a loss in speech quality, while an underestimation will lead to an unnecessary high level of residual noise. We present a novel approach for noise tracking, which updates the noise PSD for each DFT coefficient in the presence of both speech and noise. This method is based on the eigenvalue decomposition of correlation matrices that are constructed from time series of noisy DFT coefficients. The presented method is very well capable of tracking gradually changing noise types. In comparison to state-of-the-art noise tracking algorithms the proposed method reduces the estimation error between the estimated and the true noise PSD. In combination with an enhancement system the proposed method improves the segmental SNR with several decibels for gradually changing noise types. Listening experiments show that the proposed system is preferred over the state-of-the-art noise tracking algorithm.
Richard C. Hendriks, Jesper Jensen 0001, Richard Heusdens
IEEE Trans. Speech Audio Process.2
2007 DFT domain subspace based noise tracking for speech enhancement
abstract
Most DFT domain based speech enhancement methods are de-pendent on an estimate of the noise power spectral density (PSD). For non-stationary noise sources it is desirable to es-timate the noise PSD also in spectral regions where speech is present. In this paper a new method for noise tracking is pre-sented, based on eigenvalue decompositions of correlation ma-trices that are constructed from time series of noisy DFT coef-ficients. The presented method can estimate the noise PSD at time-frequency points where both speech and noise are present. In comparison to state-of-the-art noise tracking algorithms the proposed algorithm reduces the estimation error between the estimated and the true noise PSD and improves segmental SNR when combined with an enhancement system with several dB. Index Terms: Speech enhancement, noise tracking, DFT do-main subspace decompositions.
Richard C. Hendriks, Jesper Jensen 0001, Richard Heusdens
INTERSPEECH2
2007 A data-driven approach to optimizing spectral speech enhancement methods for various error criteria
Jan S. Erkelens, Jesper Jensen 0001, Richard Heusdens
Speech Commun.2
2007 Minimum Mean-Square Error Estimation of Discrete Fourier Coefficients With Generalized Gamma Priors
abstract
This paper considers techniques for single-channel speech enhancement based on the discrete Fourier transform (DFT). Specifically, we derive minimum mean-square error (MMSE) estimators of speech DFT coefficient magnitudes as well as of complex-valued DFT coefficients based on two classes of generalized gamma distributions, under an additive Gaussian noise assumption. The resulting generalized DFT magnitude estimator has as a special case the existing scheme based on a Rayleigh speech prior, while the complex DFT estimators generalize existing schemes based on Gaussian, Laplacian, and Gamma speech priors. Extensive simulation experiments with speech signals degraded by various additive noise sources verify that significant improvements are possible with the more recent estimators based on super-Gaussian priors. The increase in perceptual evaluation of speech quality (PESQ) over the noisy signals is about 0.5 points for street noise and about 1 point for white noise, nearly independent of input signal-to-noise ratio (SNR). The assumptions made for deriving the complex DFT estimators are less accurate than those for the magnitude estimators, leading to a higher maximum achievable speech quality with the magnitude estimators.
Jan S. Erkelens, Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
IEEE Trans. Speech Audio Process.4
2007 An MMSE Estimator for Speech Enhancement Under a Combined Stochastic-Deterministic Speech Model
abstract
Although many discrete Fourier transform (DFT) domain-based speech enhancement methods rely on stochastic models to derive clean speech estimators, like the Gaussian and Laplace distribution, certain speech sounds clearly show a more deterministic character. In this paper, we study the use of a deterministic model in combination with the well-known stochastic models for speech enhancement. We derive a minimum mean-square error (MMSE) estimator under a combined stochastic-deterministic speech model with speech presence uncertainty and show that for different distributions of the DFT coefficients the combined stochastic-deterministic speech model leads to improved performance of approximately 0.8 dB segmental signal-to-noise ratio (SNR) over the use of a stochastic model alone. Evaluation with perceptual evaluation of speech quality (PESQ) shows performance improvements of approximately 0.15 on an MOS scale
Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
IEEE Trans. Speech Audio Process.3
2007 Improved Subspace-Based Single-Channel Speech Enhancement Using Generalized Super-Gaussian Priors
abstract
Traditional single-channel subspace-based schemes for speech enhancement rely mostly on linear minimum mean-square error estimators, which are globally optimal only if the Karhunen-Loeacuteve transform (KLT) coefficients of the noise and speech processes are Gaussian distributed. We derive in this paper subspace-based nonlinear estimators assuming that the speech KLT coefficients are distributed according to a generalized super-Gaussian distribution which has as special cases the Laplacian and the two-sided Gamma distribution. As with the traditional linear estimators, the derived estimators are functions of the a priori signal-to-noise ratio (SNR) in the subspaces spanned by the KLT transform vectors. We propose a scheme for estimating these a priori SNRs, which is in fact a generalization of the "decision-directed" approach which is well-known from short-time Fourier transform (STFT)-based enhancement schemes. We show that the proposed a priori SNR estimation scheme leads to a significant reduction of the residual noise level, a conclusion which is confirmed in extensive objective speech quality evaluations as well as subjective tests. We also show that the derived estimators based on the super-Gaussian KLT coefficient distribution lead to improvements for different noise sources and levels as compared to when a Gaussian assumption is imposed
Jesper Jensen 0001, Richard Heusdens
IEEE Trans. Speech Audio Process.1
2007 High-Resolution Spherical Quantization of Sinusoidal Parameters
abstract
Sinusoidal coding is an often employed technique in low bit-rate audio coding. Therefore, methods for efficient quantization of sinusoidal parameters are of great importance. In this paper, we use high-resolution assumptions to derive analytical expressions for the optimal entropy-constrained unrestricted spherical quantizers for the amplitude, phase, and frequency parameters of the sinusoidal model. This is done both for the case of a single sinusoid, and for the more practically relevant case of multiple sinusoids distributed across multiple segments. To account for psychoacoustical effects of the auditory system, a perceptual distortion measure is used. The optimal quantizers minimize a high-resolution approximation of the expected perceptual distortion, while the corresponding quantization indices satisfy an entropy constraint. The quantizers turn out to be flexible and of low complexity, in the sense that they can be determined easily for varying bit rate requirements, without any sort of retraining or iterative procedures. In an objective comparison it is shown that for the squared error distortion measure, the rate-distortion performance of the proposed method is very close to that of the theoretically optimal entropy-constrained vector quantization. Furthermore, for the perceptual distortion measure, the proposed scheme is shown to objectively outperform an existing sinusoidal quantization scheme, where frequency quantization is done independently. Finally, a subjective listening test, in which the proposed scheme is compared to an existing state-of-the-art sinusoidal quantization scheme with fixed quantizers for all input signals, indicates that the proposed scheme leads to an average bit rate reduction of 20%, at the same subjective quality level as the existing scheme
Pim Korten, Jesper Jensen 0001, Richard Heusdens
IEEE Trans. Speech Audio Process.2
2006 Noise Power Spectrum Estimation for Speech Enhancement Using an Autoregressive Model for Speech Power Spectrum Dynamics
abstract
In this paper we propose a method for estimating the non-stationary noise power spectral density (PSD) given a noisy speech signal. The method is based on an autoregressive (AR) model of the speech PSD dynamics combined with a Kalman filtering based noise PSD estimation technique. Objective and subjective performance evaluations show that the speech enhancement scheme utilizing the proposed noise PSD estimation technique achieves significant improvements over a system using a stationary noise estimate as well as compared to a system that uses a noise tracker developed in our previous work
Ivo Batina, Jesper Jensen 0001, Richard Heusdens
ICASSP (3)2
2006 Speech Enhancement Under a Combined Stochastic-Deterministic Model
abstract
Most DFT domain based enhancement methods rely on stochastic models to derive clean speech estimators. In this paper we investigate the use of a deterministic speech model and present an MMSE estimator under a combined stochastic-deterministic speech model. Experimental results show an increase in segmental SNR of 1.18 dB, compared to the use of a stochastic model alone. Furthermore, PESQ evaluations lead to an increase of 0.3 on the MOS scale. Listening tests show a preference for the proposed MMSE estimator under combined stochastic-deterministic speech model
Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
ICASSP (1)3
2006 High Resolution Spherical Quantization of Sinusoids with Harmonically Related Frequencies
abstract
Sinusoidal coding is an essential tool in low-rate audio coding, and developing an efficient quantization scheme for the sinusoidal parameters is therefore crucial. In this work we derive optimal entropy constrained amplitude, phase and frequency quantizers for sinusoids whose frequencies are harmonically related, with respect to the l2distortion measure. This scheme exploits the harmonic structure of many speech and audio signals in the sense that besides amplitudes and phases, only fundamental frequencies need to be quantized, resulting in a significant decrease in the number of bits assigned to frequency parameters. The asymptotically optimal quantizers minimize a high-resolution approximation of the expected l2distortion while the corresponding quantization indices satisfy an entropy constraint. The quantizers turn out to be flexible and of low complexity, in the sense that they can be determined easily for varying bit rate requirements, without any sort of retraining or iterative procedures. In an objective rate-distortion comparison, the proposed scheme is shown to outperform two variants of a recently proposed scheme, in which all frequency parameters are quantized separately, either directly or differentially
Pim Korten, Jesper Jensen 0001, Richard Heusdens
ICASSP (5)2
2006 Perceptual Audio Coding Using N-Channel Lattice Vector Quantization
abstract
We consider the problem of reliable distribution of audio over packet-switched networks. We make use of multiple-description coding combined with transform coding in order to obtain robustness towards packet losses. Previous approaches to this problem were restricted to the case of only two descriptions. In this work we use n-channel multiple-description lattice vector quantizers (MD-LVQs), which allow for the possibility of using more than two descriptions. For a given packet-loss probability we find the number of descriptions and the bit allocation between transform coefficients which minimizes a perceptual distortion measure subject to an entropy constraint. The optimal quantizers are presented in closed form, thus avoiding any iterative quantizer design procedures. The theoretical results are verified with numerical computer simulations using audio signals and it is shown that in environments with excessive packet losses it is advantageous to use more than two descriptions. We verify in subjective listening tests that using more than two descriptions lead to signals of perceptually higher quality
Jan Østergaard, Omar Niamut, Jesper Jensen 0001, Richard Heusdens
ICASSP (5)3
2006 MMSE estimation of complex-valued discrete Fourier coefficients with generalized gamma priors
abstract
We consider DFT based techniques for single-channel speech enhancement. Specifically, we derive minimum mean-square error estimators of clean speech DFT coefficients based on generalized gamma prior probability density functions. Our estimators contain as special cases the well-known Wiener estimator and the more recently derived estimators based on Laplacian and twosided gamma priors. Simulation experiments with speech signals degraded by various additive noise sources verifythat theestimator based on the two-sided gamma prior is close to optimal amongst all the estimators considered in this paper.
Jesper Jensen 0001, Richard C. Hendriks, Jan S. Erkelens, Richard Heusdens
INTERSPEECH1
2006 Source-Channel Erasure Codes with Lattice Codebooks for Multiple Description Coding
abstract
It was recently shown that a subset of the rate distortion region of the symmetric K-channel multiple description coding problem can be achieved by use of (K, k) source-channel erasure codes (SCEC). The construction of the previously proposed SCEC made use of source coding with side information and relied upon random codebooks. In this paper we propose to replace the random codebooks of (K, k) SCEC by structured (lattice) codebooks. We then show that, in certain cases and under high-resolution assumptions, this improves the achievable rate distortion region over the traditional (K, k) SCEC
Jan Østergaard, Richard Heusdens, Jesper Jensen 0001
ISIT3
2006 Adaptive Time Segmentation for Improved Speech Enhancement
abstract
Single-channel enhancement algorithms are widely used to overcome the degradation of noisy speech signals. Speech enhancement gain functions are typically computed from two quantities, namely, an estimate of the noise power spectrum and of the noisy speech power spectrum. The variance of these power spectral estimates degrades the quality of the enhanced signal and smoothing techniques are, therefore, often used to decrease the variance. In this paper, we present a method to determine the noisy speech power spectrum based on an adaptive time segmentation. More specifically, the proposed algorithm determines for each noisy frame which of the surrounding frames should contribute to the corresponding noisy power spectral estimate. Further, we demonstrate the potential of our adaptive segmentation in both maximum likelihood and decision direction-based speech enhancement methods by making a better estimate of the a priori signal-to-noise ratio (SNR) xi. Objective and subjective experiments show that an adaptive time segmentation leads to significant performance improvements in comparison to the conventionally used fixed segmentations, particularly in transitional regions, where we observe local SNR improvements in the order of 5 dB
Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
IEEE Trans. Speech Audio Process.3
2006 Rate-distortion optimal time-segmentation and redundancy selection for VoIP
abstract
In this paper, novel techniques for packet loss robust speech coding are proposed. By exploiting knowledge of the receiving end packet loss concealment algorithm, an existing rate-distortion optimal time-segmentation algorithm is extended to taking packet losses into account. To increase robustness in highly nonstationary signals, the technique is complemented by a redundancy selection scheme. A jointly optimal approach ensures that the complementarity between time-segmentation and redundancies is fully exploited. The performance of the methods is investigated through Monte Carlo simulations under various conditions, such as rate, packet loss probability, and algorithmic delay. Finally, subjective listening tests demonstrate perceptual improvements as compared to conventional adaptive time-segmentation not taking packet losses into account.
Christoffer Rødbro, Jesper Jensen 0001, Richard Heusdens
IEEE Trans. Speech Audio Process.2
2006 n-channel entropy-constrained multiple-description lattice vector quantization
abstract
In this paper, we derive analytical expressions for the central and side quantizers which, under high-resolution assumptions, minimize the expected distortion of a symmetric multiple-description lattice vector quantization (MD-LVQ) system subject to entropy constraints on the side descriptions for given packet-loss probabilities. We consider a special case of the general n-channel symmetric multiple-description problem where only a single parameter controls the redundancy tradeoffs between the central and the side distortions. Previous work on two-channel MD-LVQ showed that the distortions of the side quantizers can be expressed through the normalized second moment of a sphere. We show here that this is also the case for three-channel MD-LVQ. Furthermore, we conjecture that this is true for the general n-channel MD-LVQ. For given source, target rate, and packet-loss probabilities we find the optimal number of descriptions and construct the MD-LVQ system that minimizes the expected distortion. We verify theoretical expressions by numerical simulations and show in a practical setup that significant performance improvements can be achieved over state-of-the-art two-channel MD-LVQ by using three-channel MD-LVQ.
Jan Østergaard, Jesper Jensen 0001, Richard Heusdens
IEEE Trans. Inf. Theory2
2005 n-Channel Symmetric Multiple-Description Lattice Vector Quantization
abstract
We derive analytical expressions for the central and side quantizers in an n-channel symmetric multiple-description lattice vector quantizer which, under high-resolution assumptions, minimize the expected distortion subject to entropy constraints on the side descriptions for given packet-loss probabilities. The performance of the central quantizer is lattice dependent whereas the performance of the side quantizers is lattice independent. In fact the normalized second moments of the side quantizers are given by that of an L-dimensional sphere. Furthermore, our analytical results reveal a simple way to determine the optimum number of descriptions. We verify theoretical results with numerical experiments and show that with a packet-loss probability of 5%, a gain of 9.1 dB in MSE over state-of-the-art two-description systems can be achieved when quantizing a two-dimensional unit-variance Gaussian source using a total bit budget of 15 bits/dimension and using three descriptions. With 20% packet loss, a similar experiment reveals an MSE reduction of 10.6 dB when using four descriptions.
Jan Østergaard, Jesper Jensen 0001, Richard Heusdens
DCC2
2005 Adaptive Time Segmentation of Noisy Speech for Improved Speech Enhancement
abstract
Enhancement algorithms are widely used to overcome the degradation of noisy speech signals. Most enhancement algorithms require an estimate of the noise and noisy speech power spectra in order to compute the gain function used for the noise suppression. The variance of these power spectral estimates degrades the quality of the enhanced signal and smoothing techniques are therefore often used to decrease the variance. We present a method to determine the noisy speech power spectrum based on an adaptive time segmentation. More specifically, the proposed algorithm determines for each noisy frame which of the surrounding frames should contribute to the corresponding noisy power spectral estimate. Objective and subjective experiments show that an adaptive time segmentation leads to significant performance improvements, particularly in transitional speech regions.
Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
ICASSP (1)3
2005 Jointly optimal time segmentation, component selection and quantization for sinusoidal coding of audio and speech
abstract
We propose a rate-distortion optimal algorithm for sinusoidal modeling of audio and speech. The algorithm determines, for a pre-specified target bit-rate, the optimal (variable-length) time segmentation, the optimal distribution of sinusoidal components over the segments and the optimal (scalar) quantizers for quantizing the sinusoid parameters. The optimization is done by jointly optimizing the segment lengths, number of sinusoids and quantizers using high-resolution quantization theory and dynamic programming techniques, which makes it possible to solve the algorithm in polynomial time. A particular advantage of the proposed method is that, given a target bit-rate, it solves the problem of finding the optimal balance between total number of sinusoids and number of bits per sinusoid.
Richard Heusdens, Jesper Jensen 0001
ICASSP (3)2
2005 High resolution spherical quantization of sinusoidal parameters using a perceptual distortion measure
abstract
Sinusoidal modelling is a key technology in low rate audio coding, and methods for efficient quantization of sinusoidal parameters are therefore of high importance. We derive analytical formulas for the optimal entropy constrained unrestricted spherical quantizers for amplitude, phase and frequency, using a perceptual distortion measure. This is done both for a single sinusoid, and for multiple sinusoids distributed over multiple segments. The quantizers minimize a high-resolution approximation of the expected distortion, while the corresponding quantization indices satisfy an entropy constraint. The quantizers turn out to be flexible and of low complexity, in the sense that they can be determined easily for varying bit rate requirements, without any sort of retraining or iterative procedures. In objective and subjective comparison tests, the proposed method is shown to outperform an existing state-of-the-art sinusoidal quantization scheme, where quantization of frequency parameters is done independently.
Pim Korten, Jesper Jensen 0001, Richard Heusdens
ICASSP (3)2
2005 Improved decision directed approach for speech enhancement using an adaptive time segmentation
abstract
Short-time Fourier transform (STFT) methods are often used to overcome the degradation of speech signals affected by noise. STFT-gain functions are usually expressed as a function of the a priori SNR, say ξ, and good techniques to estimate ξ are of vital importance for the quality of enhanced speech. Often, ξ is estimated using the so-called decision directed approach (DD). However, the DD approach builds on a number of approximations, where certain expected values of signal related quantities are approximated by instantaneous estimates. In this paper we present a method to improve these approximations by combining the DD approach with an adaptive time segmentation. Objective and subjective experiments show that the proposed method leads to significant improvements compared to the conventional DD approach. Furthermore, simulation experiments confirm a decreased amount of non-stationary residual noise.
Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
INTERSPEECH3
2005 n-channel asymmetric multiple-description lattice vector quantization
abstract
We present analytical expressions for optimal entropy-constrained multiple-description lattice vector quantizers which, under high-resolutions assumptions, minimize the expected distortion for given packet-loss probabilities. We consider the asymmetric case where packet-loss probabilities and side entropies are allowed to be unequal and find optimal quantizers for any number of descriptions in any dimension. We show that the normalized second moments of the side-quantizers are given by that of an L-dimensional sphere independent of the choice of lattices. Furthermore, we show that the optimal bit-distribution among the descriptions is not unique. In fact, within certain limits, bits can be arbitrarily distributed
Jan Østergaard, Richard Heusdens, Jesper Jensen 0001
ISIT3
2004 Perceptual linear predictive noise modelling for sinusoid-plus-noise audio coding
abstract
Sinusoidal coding of an audio subject to a bit-rate constraint, in general, results in a noise-like residual signal. This residual signal is of high perceptual importance; reconstruction of audio using the sinusoidal representation only typically results in an artificial sounding reconstruction. We present a new method, called perceptual linear predictive coding (PLPC), where the residual is encoded by applying LPC in the perceptual domain. This method minimizes a perceptual modelling error and therefore represents only residual components that are of perceptual relevance, while automatically discarding components masked by the sinusoidally coded part. Subjective listening tests show that PLPC performs significantly better than ordinary LPC as a sinusoidal residual coding technique. Furthermore, PLPC combined with a flexible segmentation and model order allocation algorithm leads to a significant gain in terms of R/D performance for fragments with fast changing characteristics.
Richard C. Hendriks, Richard Heusdens, Jesper Jensen 0001
ICASSP (4)3
2004 Entropy constrained multiple description lattice vector quantization
abstract
Recently, lattice vector quantizers (LVQ) capable of performing close to known information theoretic bounds were introduced in the area of multiple description coding (MDC). We derive analytical expressions for the central and side quantizers which minimize the expected distortion of an LVQ subject to entropy constraints on the side descriptions for given packet loss probabilities. We show that for certain packet loss probabilities, an optimal LVQ for single descriptions might not be optimal for multiple descriptions. Specifically, we show that the Z/sup 2/ lattice performs better than the A/sub 2/ lattice in some cases. Moreover, our results suggest a practical way of determining which lattice quantizers are optimal for given packet loss probabilities.
Jan Østergaard, Jesper Jensen 0001, Richard Heusdens
ICASSP (4)2
2004 Adaptive time-segmentation for speech coding with limited delay
abstract
We investigate the trade-off between delay and signal quality in adaptive time-segmentation for speech coding. A variable rate sinusoidal coder with adaptive segmentation and bit allocations is proposed and implemented with specifiable look-ahead. Objective and subjective results indicate that adaptive time-segmentation is advantageous even with low delay (30 ms), and that quality only increases with the delay until approximately 100 ms.
Christoffer Rødbro, Jesper Jensen 0001, Richard Heusdens
ICASSP (1)2
2004 A perceptual subspace approach for modeling of speech and audio signals with damped sinusoids
abstract
The problem of modeling a signal segment as a sum of exponentially damped sinusoidal components arises in many different application areas, including speech and audio processing. Often, model parameters are estimated using subspace based techniques which arrange the input signal in a structured matrix and exploit the so-called shift-invariance property related to certain vector spaces of the input matrix. A problem with this class of estimation algorithms, when used for speech and audio processing, is that the perceptual importance of the sinusoidal components is not taken into account. In this work we propose a solution to this problem. In particular, we show how to combine well-known subspace based estimation techniques with a recently developed perceptual distortion measure, in order to obtain an algorithm for extracting perceptually relevant model components. In analysis-synthesis experiments with wideband audio signals, objective and subjective evaluations show that the proposed algorithm improves perceived signal quality considerable over traditional subspace based analysis methods.
Jesper Jensen 0001, Richard Heusdens, Søren Holdt Jensen
IEEE Trans. Speech Audio Process.1
2003 A perceptual subspace method for sinusoidal speech and audio modeling
abstract
The problem of modeling a signal segment as a sum of exponentially damped sinusoidal components is of interest in a wide range of fields, including speech and audio processing. Often, model parameters are estimated using subspace based techniques that exploit the so-called shift-invariance property. A drawback of these estimation techniques in relation to speech and audio processing is that the perceptual relevance of the model components is not taken into account. In this paper we show how to combine well-known subspace based estimation techniques with a recently developed perceptual distortion measure, to obtain an algorithm for extracting perceptually relevant model components. In analysis-synthesis experiments with wideband audio signals, objective and subjective evaluations show that the proposed algorithm improves perceived signal quality considerably over traditional subspace based analysis methods.
Jesper Jensen 0001, Richard Heusdens, Søren Holdt Jensen
ICASSP (5)1
2003 Schemes for optimal frequency-differential encoding of sinusoidal model parameters
Jesper Jensen 0001, Richard Heusdens
Signal Process.1
2002 Optimal frequency-differential encoding of sinusoidal model parameters
abstract
Sinusoidal coding has proven to be efficient for low bit-rate audio coding. In this paper we consider schemes for frequency-differential (FD) encoding of the sinusoidal model parameters. For a given signal frame, the parameters of a sinusoidal component may be encoded either differentially relative to other components in the same frame, or directly, i.e., without differential encoding. Using basic tools from graph theory, two algorithms are derived for finding bit rate optimal combinations of direct and differential encoding of the sinusoidal parameters. In simulation experiments with audio signals, the algorithms showed bit-rate reductions of up to 27% relative to direct encoding. Furthermore, when compared to a commonly used FD encoding scheme, the proposed algorithms achieved bit rate reductions of up to 7%.
Jesper Jensen 0001, Richard Heusdens
ICASSP1
2001 Speech enhancement using a constrained iterative sinusoidal model
abstract
This paper presents a sinusoidal model based algorithm for enhancement of speech degraded by additive broad-band noise. In order to ensure speech-like characteristics observed in clean speech, smoothness constraints are imposed on the model parameters using a spectral envelope surface (SES) smoothing procedure. Algorithm evaluation is performed using speech signals degraded by additive white Gaussian noise. Distortion as measured by objective speech quality scores showed a 34%-41% reduction over a SNR range of 5-to-20 dB. Objective and subjective evaluations also show considerable improvement over traditional spectral subtraction and Wiener filtering based schemes. Finally, in a subjective AB preference test, where enhanced signals were coded with the G729 codec, the proposed scheme was preferred over the traditional enhancement schemes tested for SNRs in the range of 5 to 20 dB.
Jesper Jensen 0001, John H. L. Hansen
IEEE Trans. Speech Audio Process.1
2000 Harmonic exponential modeling of transitional speech segments
abstract
An extended sinusoidal model for speech signal processing is proposed. In this model, the frequencies of the sinusoidal components are constrained to be harmonically related, but their amplitudes are allowed to vary exponentially with time. Algorithms for model parameter estimation are derived. The proposed model shows considerable objective and subjective improvements in speech segments with transitions, especially in speech onsets, compared to the constant-frequency, constant-amplitude harmonic model. Potential application areas include speech synthesis and multimode speech coding.
Jesper Jensen 0001, Søren Holdt Jensen, Egon Hansen
ICASSP1
2000 Speech enhancement based on a constrained sinusoidal model
abstract
An algorithm for enhancement of speech degraded by additive broad-band noise is proposed. The algorithm represents speech using a sinusoidal model, where model parameters are estimated iteratively. In order to ensure speech-like characteristics observed in clean speech, the model parameters are restricted to satisfy certain smoothness constraints. The algorithm is evaluated using speech signals degraded by additive white Gaussian noise. Results from both objective and subjective evaluations show considerable improvement over traditional spectral subtraction and Wiener filtering based schemes. In particular, in a subjective AB preference test, where enhanced signals were encoded/decoded with the G729 speech codec, the proposed scheme was preferred over the traditional schemes in more than 5 out of 6 cases for input SNRs ranging from 5-20 dB. 1.
Jesper Jensen 0001, John H. L. Hansen
INTERSPEECH1
1999 Exponential sinusoidal modeling of transitional speech segments
abstract
A generalized sinusoidal model for speech signal processing is studied. The main feature of the model is that the amplitude of each sinusoidal component is allowed to vary exponentially with time. We propose to use the model in transitional speech segments such as speech onsets and voiced/unvoiced transitions. Computer simulations with natural speech signals indicate substantial better modeling performance in both transitional and voiced regions compared with the traditional constant-amplitude sinusoidal model.
Jesper Jensen 0001, Søren Holdt Jensen, Egon Hansen
ICASSP1